Skip to content

Size the thread pool to the container's CPU quota - #4771

Merged
springfall2008 merged 2 commits into
springfall2008:mainfrom
amasolov:container-aware-cpu-count
Aug 27, 2026
Merged

springfall2008 merged 2 commits into
springfall2008:mainfrom
amasolov:container-aware-cpu-count

Conversation

@amasolov

Copy link
Copy Markdown
Contributor

The problem

threads: auto sizes the batch pool from cpu_count(), which reports the machine's cores. Inside a container that is the wrong number: a Docker --cpus or a Kubernetes CPU limit is a cgroup bandwidth quota, and the host's full core count stays visible through both cpu_count() and sched_getaffinity(). Predbat sizes its pool to the host, and the CFS scheduler then throttles it.

I hit this running Predbat on Kubernetes with a 4-core limit on a 12-core node:

os.cpu_count()          -> 12
sched_getaffinity        -> 12
/sys/fs/cgroup/cpu.max   -> 400000 100000   # i.e. 4 cores

The pod sat pegged at its quota, CPU-throttled, for the whole of a 30-epoch ML fine-tune — twelve lanes contending for four cores, which is slower than four lanes would have been. Note sched_getaffinity() is also 12, so the usual affinity-based workaround does not help here; it has to be the quota.

The change

available_cpu_count() reads the cgroup quota (v2 cpu.max, then v1 cpu.cfs_quota_us/cpu.cfs_period_us) and falls back to cpu_count() on bare metal, in a container with no limit set, and on any platform without cgroups. Behaviour outside a constrained container is unchanged.

This deliberately does not touch the "auto is not capped" policy. The reasoning in resolve_batch_threads's docstring — that a cap costs 10.7% on a kernel-heavy machine against 1.3% for no cap — is about physical cores, and I read it as still correct. This only changes which number auto is uncapped against: the CPUs the process may actually use, rather than the ones it can see. A quota is a hard ceiling the scheduler enforces regardless, so exceeding it buys contention rather than throughput.

Rounding is down (a 3.5-core quota sustains three fully-busy lanes), never below 1, and never above the host count.

Tests

Added to test_prediction_batch.py alongside the existing thread-resolution test, and registered in run_prediction_batch_tests so a regression fails the suite.

The cgroup reads are stubbed rather than exercised against the real filesystem, so the test gives the same answer on a laptop, inside a container, and in CI. Cases cover v2, v1, unlimited in both, fractional and sub-core quotas, a quota above the host count, no cgroup files at all, and an unparsable quota.

I checked the test fails against the unfixed code rather than only passing against the new code — reverting the quota branch produces four ERROR: lines and failed = True.

black --line-length 256 clean; flake8 --max-line-length 250 reports nothing new (the remaining F841/C901 hits are pre-existing on main).

Happy to adjust the approach — including capping auto at the quota only when one is present and leaving everything else alone, if you would rather keep the change narrower.

threads: auto sizes the batch pool from cpu_count(), which reports the
machine's cores. Inside a container that is the wrong number: a Docker
--cpus or a Kubernetes CPU limit is a cgroup bandwidth quota, and the
host's full core count stays visible through both cpu_count() and
sched_getaffinity(). Predbat therefore sizes its pool to the host and the
CFS scheduler throttles it.

Seen on a Kubernetes pod limited to 4 cores on a 12-core node:
cpu_count() reports 12 while /sys/fs/cgroup/cpu.max reports
"400000 100000". The pod sat pegged at its quota, CPU-throttled, for the
whole of a 30-epoch ML fine-tune - twelve lanes contending for four
cores, which is slower than four lanes would have been.

available_cpu_count() reads the cgroup quota (v2 and v1) and falls back
to cpu_count() on bare metal, in a container with no limit set, and on
any platform without cgroups - so nothing changes outside a constrained
container. The 'auto is not capped' reasoning in resolve_batch_threads is
untouched: this changes which number 'auto' is uncapped against, not the
policy.

Tests stub the cgroup reads rather than touching the real filesystem, so
they give the same answer on a laptop, in a container and in CI. Verified
they fail against the unfixed code rather than only passing against the
new code.
The docstring on available_cpu_count() names sched_getaffinity(), since
its returning the host's core count is why the cgroup quota has to be
read instead. cspell does not know the word - it sits next to getrusage,
which is already in the dictionary.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The quota-aware CPU sizing is implemented with conservative fallbacks and is covered by focused, deterministic unit tests for the key cgroup scenarios.

Pull request overview

This PR improves Predbat’s threads: auto sizing inside containers by basing the kernel batch thread count on the container’s effective CPU quota (cgroup v2/v1) rather than the host’s visible core count, avoiding excessive contention and CFS throttling under CPU limits.

Changes:

  • Add available_cpu_count() to read cgroup CPU quotas (v2 cpu.max, then v1 cpu.cfs_quota_us/cpu.cfs_period_us) with safe fallbacks to the host core count.
  • Use available_cpu_count() when resolving self.prediction.batch_threads for batched prediction kernel execution.
  • Add unit tests covering v2/v1 quotas, unlimited cases, fractional/sub-core quotas, missing files, and parse failures; update cspell dictionary for “getaffinity”.
File summaries
File Description
apps/predbat/plan.py Introduces quota-aware CPU availability and uses it to size the prediction kernel thread lanes under threads: auto.
apps/predbat/tests/test_prediction_batch.py Adds deterministic tests that stub cgroup reads to validate CPU quota parsing and fallback behavior.
.cspell/custom-dictionary-workspace.txt Adds “getaffinity” to avoid spellcheck noise from the updated documentation text.
Review details
  • Files reviewed: 3/3 changed files
  • Comments generated: 0
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@springfall2008
springfall2008 merged commit fa68a92 into springfall2008:main Aug 27, 2026
2 checks passed
@amasolov
amasolov deleted the container-aware-cpu-count branch August 28, 2026 02:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants