Skip to content

Composite tasks do not notify parent composites #266

Description

WhenAllTask and WhenAnyTask mark themselves complete, but do not notify their parent composite.

This means the following orchestration does not resume when all activities complete:

all_task = task.when_all(activity_tasks)
winner = yield task.when_any([cancel_task, all_task])

A current workaround is to wait for each leaf task with repeated when_any calls. The workaround preserves cancellation, but Durable Task Scheduler no longer shows the standard Fan-out group.

Could composite completion propagate to the parent in the same way that CompletableTask.complete() does?

A regression test can compose when_any([cancel, when_all([a, b])]), complete a and b, and verify that the outer task completes with the inner when_all task.

Compatibility and behavior-change considerations

PR #267 fixes the missing completion propagation unconditionally. The proposed
direction is to accept this intentional replay-compatibility break rather than
preserve the bug, provided the affected scenarios and upgrade guidance are
explicit. This is a behavioral/replay change, not a public signature change or
a new abstract-method requirement for third-party subclasses.

Affected histories

An existing instance can have a history like this:

  1. when_any([cancel, when_all(tasks)]) is pending.
  2. All activities finish, but the inner when_all does not notify the race.
  3. A later cancellation wins, and the orchestration records cancellation-branch
    work, for example a cleanup activity.
  4. After upgrading, replay completes the inner composite at step 2 and takes the
    success branch instead. If that branch schedules finish where history
    recorded cleanup, replay fails with NonDeterminismError.

Warning

A changed replay is not guaranteed to raise NonDeterminismError. If both
branches schedule the same activity name/ID but use different inputs or local
continuation state, the executor can accept the recorded action and continue
with the new branch. An old history using record("cleanup") was reproduced
completing with "cleanup" on the merge base and "finish" with the fix.
Do not rely on nondeterminism errors alone to identify affected instances.

This is not a blanket break for all orchestrations. Ordinary non-nested
composites retain their behavior. A cancellation-first history for the example
above still replays consistently. A history still blocked at the nested
composite, with no conflicting downstream continuation, can regain progress
when the orchestration next executes. Deploying the fix alone does not
necessarily wake an otherwise idle instance.

Provider and deployment implications

Both durabletask.azuremanaged and azure-functions-durable v2 use the core
implementation, including the Functions compatibility APIs context.task_all()
and context.task_any(). Their current core constraints are open-ended:
durabletask>=1.10.1 and durabletask[opentelemetry]>=1.10.1, respectively.
Pinning only the provider therefore does not freeze this behavior; a fresh
dependency resolution can pick up the correction without a provider-version
change. A core major-version bump alone would not protect those existing
provider constraints.

  • Pin or lock the core durabletask version until the transition is deliberate.
  • Let affected running instances finish on the previous SDK, or retain the old
    deployment/task hub and use a separate task hub for new instances.
  • Avoid mixing old and corrected workers on the same affected task hub during
    the transition. Stuck instances may require deliberate recovery rather than
    simply waiting to drain.
  • Keep the warnings in all three package changelogs and carry them into release
    notes, including upgrades from earlier Functions v2 prereleases. Coordinate
    package/dependency version changes in the release PR, not this fix.

PR scope and regression coverage

The latest PR already includes replay warnings in all three changelogs. The
implementation preserves result ordering, waits for all when_all children
before surfacing failure, and notifies a when_any parent only for the first
winner. No additional public API or minimum-Python-version change is needed.

The new tests exercise task objects directly. An executor-level regression
would additionally preserve the intended scheduling/replay boundary: successful
nested completion, replay of a history written by the corrected implementation,
and an unchanged cancellation-first history. The legacy-history mismatch above
is also useful to encode as an intentional compatibility test.

The automated review's empty-when_all([]) concern reproduces on both the merge
base and the PR. It is a pre-existing empty-input behavior, not a regression
introduced by completion propagation. Keep any empty-input semantic change
separate rather than silently expanding this fix.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions