Huawei · Ascend Training Development
AI Training, Inference & Heterogeneous Acceleration Engineer
- Optimized PyTorch / torch_npu training paths on Ascend NPUs, including fine-grained CPU affinity, operator Lazy Init, parameter tuning, and L2-cache utilization.
- Adapted OpenSoraPlan, DeepSeek-VL2, and GLM-4.1V with Megatron / FSDP; contributed to DeepSeek-R1 and Qwen3 large-EP inference and Attention INT8 quantization.
- Adapted UniVL and verl + Qwen2.5-VL multimodal RL pipelines; adapted DanceGRPO and raised a generative-model RL score from 0.4 to 0.8.
- Doubled Intern-S1 Pro inference through efficient GMM and post-processing computation; optimized GDN Triton kernels and the framework to improve Intern-S2 Preview (Qwen3.5-35B) SFT from 0.12× to 0.72× H200.
