feat:Optimize qwen2-vl to reduce cudaMemcpyAsync #14377

cynthieye · 2025-03-06T17:51:47Z

qwen2-vl logic optimization: During each forward propagation, the xformer branch of Qwen2VisionTransformer will execute multiple tensor tolist methods (flash attn branch will execute multiple tensor items) to force the GPU tensor to be copied to the CPU, triggering CUDAMemcpyAsync to increase time consumption. Since the input and output are the same multiple times, it will be executed once, and the remaining will reuse the first result

BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing/overview.html

feat:Optimize qwen2-vl to reduce cudaMemcpyAsync

3f21ec2

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

feat:Optimize qwen2-vl to reduce cudaMemcpyAsync #14377

feat:Optimize qwen2-vl to reduce cudaMemcpyAsync #14377

cynthieye commented Mar 6, 2025

feat:Optimize qwen2-vl to reduce cudaMemcpyAsync #14377

Are you sure you want to change the base?

feat:Optimize qwen2-vl to reduce cudaMemcpyAsync #14377

Conversation

cynthieye commented Mar 6, 2025