Size the thread pool to the container's CPU quota - #4771
Merged
springfall2008 merged 2 commits intoAug 27, 2026
Merged
Conversation
threads: auto sizes the batch pool from cpu_count(), which reports the machine's cores. Inside a container that is the wrong number: a Docker --cpus or a Kubernetes CPU limit is a cgroup bandwidth quota, and the host's full core count stays visible through both cpu_count() and sched_getaffinity(). Predbat therefore sizes its pool to the host and the CFS scheduler throttles it. Seen on a Kubernetes pod limited to 4 cores on a 12-core node: cpu_count() reports 12 while /sys/fs/cgroup/cpu.max reports "400000 100000". The pod sat pegged at its quota, CPU-throttled, for the whole of a 30-epoch ML fine-tune - twelve lanes contending for four cores, which is slower than four lanes would have been. available_cpu_count() reads the cgroup quota (v2 and v1) and falls back to cpu_count() on bare metal, in a container with no limit set, and on any platform without cgroups - so nothing changes outside a constrained container. The 'auto is not capped' reasoning in resolve_batch_threads is untouched: this changes which number 'auto' is uncapped against, not the policy. Tests stub the cgroup reads rather than touching the real filesystem, so they give the same answer on a laptop, in a container and in CI. Verified they fail against the unfixed code rather than only passing against the new code.
The docstring on available_cpu_count() names sched_getaffinity(), since its returning the host's core count is why the cgroup quota has to be read instead. cspell does not know the word - it sits next to getrusage, which is already in the dictionary.
Contributor
There was a problem hiding this comment.
🟢 Approval recommended
The quota-aware CPU sizing is implemented with conservative fallbacks and is covered by focused, deterministic unit tests for the key cgroup scenarios.
Pull request overview
This PR improves Predbat’s threads: auto sizing inside containers by basing the kernel batch thread count on the container’s effective CPU quota (cgroup v2/v1) rather than the host’s visible core count, avoiding excessive contention and CFS throttling under CPU limits.
Changes:
- Add
available_cpu_count()to read cgroup CPU quotas (v2cpu.max, then v1cpu.cfs_quota_us/cpu.cfs_period_us) with safe fallbacks to the host core count. - Use
available_cpu_count()when resolvingself.prediction.batch_threadsfor batched prediction kernel execution. - Add unit tests covering v2/v1 quotas, unlimited cases, fractional/sub-core quotas, missing files, and parse failures; update cspell dictionary for “getaffinity”.
File summaries
| File | Description |
|---|---|
| apps/predbat/plan.py | Introduces quota-aware CPU availability and uses it to size the prediction kernel thread lanes under threads: auto. |
| apps/predbat/tests/test_prediction_batch.py | Adds deterministic tests that stub cgroup reads to validate CPU quota parsing and fallback behavior. |
| .cspell/custom-dictionary-workspace.txt | Adds “getaffinity” to avoid spellcheck noise from the updated documentation text. |
Review details
- Files reviewed: 3/3 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The problem
threads: autosizes the batch pool fromcpu_count(), which reports the machine's cores. Inside a container that is the wrong number: a Docker--cpusor a Kubernetes CPU limit is a cgroup bandwidth quota, and the host's full core count stays visible through bothcpu_count()andsched_getaffinity(). Predbat sizes its pool to the host, and the CFS scheduler then throttles it.I hit this running Predbat on Kubernetes with a 4-core limit on a 12-core node:
The pod sat pegged at its quota, CPU-throttled, for the whole of a 30-epoch ML fine-tune — twelve lanes contending for four cores, which is slower than four lanes would have been. Note
sched_getaffinity()is also 12, so the usual affinity-based workaround does not help here; it has to be the quota.The change
available_cpu_count()reads the cgroup quota (v2cpu.max, then v1cpu.cfs_quota_us/cpu.cfs_period_us) and falls back tocpu_count()on bare metal, in a container with no limit set, and on any platform without cgroups. Behaviour outside a constrained container is unchanged.This deliberately does not touch the "auto is not capped" policy. The reasoning in
resolve_batch_threads's docstring — that a cap costs 10.7% on a kernel-heavy machine against 1.3% for no cap — is about physical cores, and I read it as still correct. This only changes which numberautois uncapped against: the CPUs the process may actually use, rather than the ones it can see. A quota is a hard ceiling the scheduler enforces regardless, so exceeding it buys contention rather than throughput.Rounding is down (a 3.5-core quota sustains three fully-busy lanes), never below 1, and never above the host count.
Tests
Added to
test_prediction_batch.pyalongside the existing thread-resolution test, and registered inrun_prediction_batch_testsso a regression fails the suite.The cgroup reads are stubbed rather than exercised against the real filesystem, so the test gives the same answer on a laptop, inside a container, and in CI. Cases cover v2, v1, unlimited in both, fractional and sub-core quotas, a quota above the host count, no cgroup files at all, and an unparsable quota.
I checked the test fails against the unfixed code rather than only passing against the new code — reverting the quota branch produces four
ERROR:lines andfailed = True.black --line-length 256clean;flake8 --max-line-length 250reports nothing new (the remainingF841/C901hits are pre-existing onmain).Happy to adjust the approach — including capping
autoat the quota only when one is present and leaving everything else alone, if you would rather keep the change narrower.