Skip to content

feat: support EPLB for DeepSeek-V3.2 Python model executor. - #2283

Draft
yingxudeng wants to merge 1 commit into
xLLM-AI:mainfrom
yingxudeng:feat/deepseek-v32-python-eplb
Draft

feat: support EPLB for DeepSeek-V3.2 Python model executor.#2283
yingxudeng wants to merge 1 commit into
xLLM-AI:mainfrom
yingxudeng:feat/deepseek-v32-python-eplb

Conversation

@yingxudeng

Copy link
Copy Markdown
Collaborator

Add Expert Parallel Load Balancing (EPLB) support to the Python MoE path. The implementation uses slot-reuse mode where device_experts_num stays at num_local_experts to avoid exceeding the NPU fused GMM kernel groupList length limit. EPLB replaces cold expert weights in-place within fixed slots rather than allocating additional redundant slots.

Key changes:

  • New xllm/python/layers/eplb.py with helper functions ported from C++
  • PyCausalLM bridges prepare/start_transfer/update/last_ok to Python
  • DeepseekV3MoE: log2phy_map remap in forward, dynamic lifecycle
  • grouped_moe kernel: log2phy_map parameter for expert id remapping

Verified EP=2 (redundant=0) and EP=4 (redundant=1) outputs match baseline (EPLB off) token-for-token on 6-layer w8a8 model.

Description

Related Issues

Change Type

  • Bug fix
  • New feature
  • Performance improvement
  • Refactor
  • Documentation
  • Test
  • Build or CI

Pull Request Checklist

Thank you for contributing to xLLM. Before requesting review, please make sure the following items are complete.

PR Title and Commit Messages

  • The PR title and each commit message follow the xLLM commit format: <type>: <subject>.

Allowed types: feat, bugfix, docs, test, refactor, chore, style, revert, perf, model, build, release.
The subject should use clear English, start with a verb, include at least 4 words, and end with ..

Pre-commit Checks

  • I have installed pre-commit by running pip install pre-commit or an equivalent command.
  • I have installed the hooks with pre-commit install.
  • I have run pre-commit run --all-files and fixed any reported issues.

If you are unsure how to set up pre-commit, see the pre-commit documentation.

Self Review

  • I have self-reviewed the code according to .agents/skills/code-review/references/custom-code-style.md, especially code written or assisted by AI.
  • I have rebased this PR onto the latest main branch.

Build and Test Coverage

  • Tests have been added or updated as needed.
  • CUDA: python setup.py build test has passed on a CUDA machine.
  • NPU: python setup.py build test has passed on an NPU machine.
  • MLU: python setup.py build test has passed on an MLU machine.

Reviewer Notes

Add Expert Parallel Load Balancing (EPLB) support to the Python MoE
path. The implementation uses slot-reuse mode where device_experts_num
stays at num_local_experts to avoid exceeding the NPU fused GMM kernel
groupList length limit. EPLB replaces cold expert weights in-place
within fixed slots rather than allocating additional redundant slots.

Key changes:
- New xllm/python/layers/eplb.py with helper functions ported from C++
- PyCausalLM bridges prepare/start_transfer/update/last_ok to Python
- DeepseekV3MoE: log2phy_map remap in forward, dynamic lifecycle
- grouped_moe kernel: log2phy_map parameter for expert id remapping

Verified EP=2 (redundant=0) and EP=4 (redundant=1) outputs match
baseline (EPLB off) token-for-token on 6-layer w8a8 model.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant