Skip to content

Ship the quantized kernels as their own library - #21642

Merged
shoumikhin merged 238 commits into
mainfrom
gh/shoumikhin/93/head
Aug 20, 2026
Merged

Ship the quantized kernels as their own library#21642
shoumikhin merged 238 commits into
mainfrom
gh/shoumikhin/93/head

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

A quantized model uses smaller numbers than a normal one, so the tensors take less memory.
Running one needs the quantized operator kernels.

The only copy the wheel shipped is the one torch loads to export a model, which a C++ application
cannot use. Such an application links the runtime, loads a quantized model, and the model fails at
run time with a missing operator, which looks like a model problem rather than a packaging one.

Build the quantized kernels as their own shared library and name it as a CMake component, the same
way the other kernel sets are named.

find_package(executorch REQUIRED COMPONENTS kernels_quantized)
target_link_libraries(my_app PRIVATE executorch::runtime
                                     executorch::kernels_quantized)

The wheel now ships lib/libexecutorch_kernels_quantized.so.

The wheel used to ship these kernels twice. This build compiled them once for the runtime
component and again for the export plugin torch loads: gen_selected_ops ran against the same yaml
under two library names, and the same source list was compiled into two targets whose only
difference was portable_lib against executorch_core on the link line. Measured on the shipped
artifacts, neither library carried an operator the other lacked, so it was a
copy rather than an overlap.

Registration happens in a static initializer, so both copies fired on load and the second aborted
the process on a repeat registration. No load order avoided it.

Now the export plugin links the shared quantized library instead of carrying its own copy, so one
registrar reaches one registry. The duplicate pair is built only where there is no shared library to
link, which is the configuration that never had two copies to begin with. The plugin already routed
kernels/quantized/ to lib/ for its runtime search path, which is where that library ships, and
the retention option that keeps a registration-only library on the link line is an interface
property, so it reaches the plugin as well. The wheel also gets smaller: the plugin was 733 KB
carrying a copy of a 303 KB library.

runtime/kernel/operator_registry.cpp is untouched. Making an identical re-registration idempotent
there, or guarding the generated registration with registry_has_op_function, would weaken a check
that catches genuinely conflicting implementations for every embedder in order to paper over one
build's duplicate, and would make the winner depend on load order.

Built the wheel, installed it into a clean environment, and:

  • exported a quantized model and ran it from Python, matching eager PyTorch to within the
    quantization step (measured worst difference 0.0048 against a tolerance of 0.02).
  • built a C++ application that links executorch::kernels_quantized, ran the same program, and got
    the same output as Python, byte for byte.
  • confirmed the Python extension does not depend on the run-time copy, and that a process holding
    the shipped library and the export plugin no longer aborts, where before this change it did in
    either load order.
  • checked every shipped library the same way, to establish that this is the only pair that
    collides: the CPU kernels, the delegate, the thread pool, the profiler and the runtime all
    coexist with both the extension and the export plugin.
  • an application linking only EXECUTORCH_LIBRARIES does not depend on the quantized library while
    still depending on the CPU kernels, on CMake 3.28 and on real CMake 3.24. A new check asserts
    this, and it fails on the previous behaviour.
  • EXECUTORCH_QUANTIZED_KERNELS_LIBRARY resolves to the shipped library on both the modern-CMake
    route (as the imported target) and the pre-3.28 route (as a file path).
  • a missing quantized library now fails the checks instead of skipping them. The preset that builds
    the wheel enables these kernels unconditionally, so their absence is a regression rather than a
    configuration to tolerate, and both the ownership table and the C++ check previously treated it as
    an acceptable state and reported coverage they had not run.

Ran on Linux x86_64 and aarch64.

@pytorch-bot

pytorch-bot Bot commented Aug 7, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21642

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 441 Pending

As of commit 1149045 with merge base 558473d (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 7, 2026
@shoumikhin shoumikhin added ciflow/periodic ciflow/trunk ciflow/binaries ciflow/binaries/all Release PRs with this label will build wheels for all python versions ciflow/nightly ciflow/cuda labels Aug 7, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/binaries/all Release PRs with this label will build wheels for all python versions ciflow/binaries ciflow/cuda ciflow/nightly ciflow/periodic ciflow/trunk CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants