Repository navigation
feat: add optimized Kunlun rotary embedding backend - #1001
Draft
JoeZhang-0x000 wants to merge 1 commit into
Draft
JoeZhang-0x000 wants to merge 1 commit into
JoeZhang-0x000 wants to merge 1 commit into
Conversation
This was referenced Oct 10, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
RotaryEmbeddingbackend and only its required build/device/Python-dispatch plumbing.Motivation
Depends on InfiniRT #48, revision
e02d72ec75414d234cb1473994f6463856844acc. Build against that runtime until it merges. Supersedes the closed InfiniCore #1574: current operator implementations belong in InfiniOps. There is no existing Kunlun RoPE on upstream master, so this PR includes the minimal backend entry points as well as the optimization.Type of Change
feat— new platform/operator backendPlatforms Affected
WITH_KUNLUN)Smoke Test Result
P800 OAM (12 clusters), XTDK LLVM 15/xpu3, XRE runtime query 5.0 (
bded6e7), GCC 9.4, PyTorch 2.5.1. Fresh native library and Python binding builds:For direct CMake installs, expose the extension as
infini.ops(the wheel already uses that namespace). Tests include 18 mixed-dtype/NeoX/GPT-J graph cases with two changed-input/position replays each: negatives,2**40,INT64_MAX, 512 rotary dimensions, offset/inverse, strided QKV and padded cache rows. All 24 BF16 benchmark shapes also match the CPU reference with zero eager error; baseline/candidate eager and changed-input graph output hashes match in every shape.Test Results on Supported Platforms
add,mul,relu,cast)clang-format 21, Ruff andgit diff --checkpassed. Full unrelated-operator suite not run; Kunlun currently implements RoPE only.Benchmark / Performance Impact
Controlled comparison on one P800, BF16 packed QKV views, Qwen3-8B head dimension 128 and Q/KV heads
32/TP,8/TP. TP labels describe per-rank shapes, not multi-card execution. Baseline is the same modern implementation with only the former 8-cluster/contiguous-worker scheme restored; candidate uses 12 clusters/striped workers. Same runtime/toolchain and native int64 API on both sides.Median synchronized graph replay time per RoPE Q+K call (μs): 8 calls/graph, 20 replays/sample, 5 samples, 3 warm-up replays. Graph execution is checked after changing inputs and positions before timing.
8192-token prefill:
Across 24 shapes (TP 1/2/4/8 × 1/4/16/64/512/8192 tokens), speedup is 0.957–1.489×. Six tiny 1–4-token cases regress by at most 0.56 μs (4.5%); larger shapes improve.
Notes for Reviewers
Kernel launch metadata uses fixed-width types because XTDK's device
size_t/ptrdiff_tdiffer from the host ABI. The public operator API is unchanged; the legacy capability symbol is unnecessary here. No attention routing or other kernels are added, and no report documents, CSV files or benchmark scripts are committed.Earlier legacy vLLM integration results, not rerun for this new API port: Qwen3-0.6B GSM8K native/plugin both 715/1319 (54.21%, gap 0 pp); Qwen3-8B static throughput minima versus the same vendor baseline at TP1/2/4/8: 111.12%/104.65%/98.47%/90.17%. Both sides used the same isolated vendor cache fix, with vendor attention. Original integration report.