Updated from the experience profile and Google Scholar

Download English PDF

AI SYSTEMS / ASCEND / HPC

Shaojie Tan

AI Systems · Ascend Training & Inference · HPC

Shaojie Tan

Huawei Ascend training systems engineer with an M.S. in Computer Science from USTC. Focused on large-model training and inference optimization on Ascend NPUs, multimodal migration, distributed parallelism, reinforcement learning, and quantization, backed by computer-architecture and multi-node CPU/GPU/NPU optimization experience.

Intern-S2 relative H200 performance
Intern-S1 Pro inference performance
3 papers HPCC · DATE · CCF Trans. HPC

Education

University of Science and Technology of ChinaM.S. in Computer Science · Computer Architecture
University of Science and Technology of ChinaB.S. in Computer Science

Experience & Research

Huawei · Ascend Training Development

AI Training, Inference & Heterogeneous Acceleration Engineer

  • Optimized PyTorch / torch_npu training paths on Ascend NPUs, including fine-grained CPU affinity, operator Lazy Init, parameter tuning, and L2-cache utilization.
  • Adapted OpenSoraPlan, DeepSeek-VL2, and GLM-4.1V with Megatron / FSDP; contributed to DeepSeek-R1 and Qwen3 large-EP inference and Attention INT8 quantization.
  • Adapted UniVL and verl + Qwen2.5-VL multimodal RL pipelines; adapted DanceGRPO and raised a generative-model RL score from 0.4 to 0.8.
  • Doubled Intern-S1 Pro inference through efficient GMM and post-processing computation; optimized GDN Triton kernels and the framework to improve Intern-S2 Preview (Qwen3.5-35B) SFT from 0.12× to 0.72× H200.
XTuner + vLLM Technical ReportQwen3.5 Optimization Report

Huawei 2012 Laboratories

Research Intern · PIM Task Scheduling

  • Built a compile-time performance model and reduced data movement through locality and connectivity clustering.
  • Selected execution between PIM and CPU using static features such as parallelism and memory-port pressure.
  • Achieved near-optimal scheduling using static analysis alone, eliminating dynamic measurement overhead.

USTC × Huawei

Kunpeng 920 Basic-Block Throughput Modeling

  • Developed a throughput prediction framework using instruction throughput, latency, and port-pressure data.
  • Built a benchmark suite spanning real applications and evaluation tools, plus a runtime measurement system.
  • Improved llvm-mca accuracy on AArch64 by modeling Kunpeng 920 microarchitecture; published at HPCC 2022.

Selected Projects

Intern-S2 Preview SFT Optimization

Pujiang Lab · Qwen3.5-35B

  • Profiled GDN at over 70% of out-of-box runtime, eliminated transpose overhead in causal_conv1d / KKT paths, and replaced or optimized key inefficient Triton implementations with Ascend C.
  • Added GDN chunk-size 128 support and optimized solve-tril, KKT, RMSNormGated, and CV-core synchronization; tuned CPU affinity, removed two-node SP2 communication, fused RMSNorm, and upgraded CANN.
  • Combined kernel and framework optimization improved Intern-S2 Preview (Qwen3.5-35B) SFT from 0.12× to 0.72× H200.
Read the Qwen3.5 Optimization Report (PDF)

Multi-GPU Quantum Circuit Simulation with QuEST

ASC20-21 Student Supercomputer Challenge

  • Enabled direct GPU-memory access with CUDA-aware MPI and overlapped computation and communication using streams and buffers.
  • Replaced vector dot products and reductions with cuBLAS, scaling from 30 to 34 qubits on 8 NVIDIA A100 GPUs across 4 nodes.
  • Saved more than 200 GB of memory versus the CPU system and won First Prize in the finals.

Multi-node CPU Optimization of a Serial Data Application

IPCC 2022 Parallel Computing Challenge

  • Combined loop unrolling, manual vectorization, OpenMP, and MPI for instruction-, node-, and cluster-level parallelism.
  • Improved data layout, bandwidth utilization, communication volume, mixed precision, and PSRS-based hotspots.
  • Reached 10,000× speedup on dual AMD EPYC 7452 nodes and up to 2,800,000× on a single node.

Technical Skills

ProgrammingC, C++, Python, Shell
AI FrameworksPyTorch, torch_npu, Megatron, FSDP, verl, vLLM
LLM SystemsAscend NPU training/inference, multimodal models, RL post-training, Expert Parallel, Triton, Ascend C, Attention INT8 quantization
Parallel ComputingMPI, OpenMP, CUDA, cuBLAS, multi-node CPU/GPU/NPU optimization
ArchitectureIntel/ARM ISA, AArch64, CPU microarchitecture, PIM, performance modeling
PerformanceAssembly analysis, vectorization, cache and bandwidth optimization, llvm-mca

Awards

Outstanding Graduate, USTC2024
ASC20-21 Student Supercomputer ChallengeFirst Prize, Finals
ISC21 Student Cluster CompetitionFirst Prize, Finals
IPCC 2022 Parallel Computing ChallengeNational Third Prize · 5/130+

Publications

S. Tan, Q. Jiang, Z. Cao, X. Hao, J. Chen, H. An. Uncovering the Performance Bottleneck of Modern HPC Processor with Static Code Analyzer: A Case Study on Kunpeng 920 CCF Transactions on High Performance Computing 6(3), 2024

Q. Jiang, S. Tan, J. Chen, H. An. A³PIM: An Automated, Analytic and Accurate Processing-in-Memory Offloader DATE 2024

Q. Jiang, S. Tan, Z. Cao, X. Hao, J. Chen, H. An. Quantifying Throughput of Basic Blocks on ARM Microarchitectures by Static Code Analyzers: A Case Study on Kunpeng 920 IEEE HPCC 2022

AI Practice & Engineering Automation

I embrace AI through hands-on projects that accelerate technical understanding and development automation: mapping technology evolution, tracking delivery, organizing stateful agent teams, and exploring performance-model-driven optimization.

AIHisTrendFuture

Live

An evidence-backed timeline of AI history, trends, and future, preserving primary sources, derivation chains, and confidence levels.

GitHubLive Site

DevTracking

In use

Turns plans, progress, blockers, demands, and delivery evidence into traceable workflows and visual dashboards with durable project context.

GitHub

IPD-Teams

Experimental

Applies IPD-style roles to stateful software-delivery agent teams, exploring multi-agent collaboration, handoffs, and quality loops.

GitHub

AutoResearch

In development

Builds a local-first LLM training and debugging loop that turns experiments, metrics, design tradeoffs, and retrospectives into reusable assets.

GitHub

ProgramModeling

In development

Builds executable graphs for memory, performance, MFU, and parallelism, progressing toward configuration recommendation, agent verification, and calibration.

GitHub

Career Vision

See the essence through first principles; amplify the value of intelligence through systems innovation.