You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A current workaround is to wait for each leaf task with repeated when_any calls. The workaround preserves cancellation, but Durable Task Scheduler no longer shows the standard Fan-out group.
Could composite completion propagate to the parent in the same way that CompletableTask.complete() does?
A regression test can compose when_any([cancel, when_all([a, b])]), complete a and b, and verify that the outer task completes with the inner when_all task.
Compatibility and behavior-change considerations
PR #267 fixes the missing completion propagation unconditionally. The proposed
direction is to accept this intentional replay-compatibility break rather than
preserve the bug, provided the affected scenarios and upgrade guidance are
explicit. This is a behavioral/replay change, not a public signature change or
a new abstract-method requirement for third-party subclasses.
Affected histories
An existing instance can have a history like this:
when_any([cancel, when_all(tasks)]) is pending.
All activities finish, but the inner when_all does not notify the race.
A later cancellation wins, and the orchestration records cancellation-branch
work, for example a cleanup activity.
After upgrading, replay completes the inner composite at step 2 and takes the
success branch instead. If that branch schedules finish where history
recorded cleanup, replay fails with NonDeterminismError.
Warning
A changed replay is not guaranteed to raise NonDeterminismError. If both
branches schedule the same activity name/ID but use different inputs or local
continuation state, the executor can accept the recorded action and continue
with the new branch. An old history using record("cleanup") was reproduced
completing with "cleanup" on the merge base and "finish" with the fix.
Do not rely on nondeterminism errors alone to identify affected instances.
This is not a blanket break for all orchestrations. Ordinary non-nested
composites retain their behavior. A cancellation-first history for the example
above still replays consistently. A history still blocked at the nested
composite, with no conflicting downstream continuation, can regain progress
when the orchestration next executes. Deploying the fix alone does not
necessarily wake an otherwise idle instance.
Provider and deployment implications
Both durabletask.azuremanaged and azure-functions-durable v2 use the core
implementation, including the Functions compatibility APIs context.task_all()
and context.task_any(). Their current core constraints are open-ended: durabletask>=1.10.1 and durabletask[opentelemetry]>=1.10.1, respectively.
Pinning only the provider therefore does not freeze this behavior; a fresh
dependency resolution can pick up the correction without a provider-version
change. A core major-version bump alone would not protect those existing
provider constraints.
Pin or lock the core durabletask version until the transition is deliberate.
Let affected running instances finish on the previous SDK, or retain the old
deployment/task hub and use a separate task hub for new instances.
Avoid mixing old and corrected workers on the same affected task hub during
the transition. Stuck instances may require deliberate recovery rather than
simply waiting to drain.
Keep the warnings in all three package changelogs and carry them into release
notes, including upgrades from earlier Functions v2 prereleases. Coordinate
package/dependency version changes in the release PR, not this fix.
PR scope and regression coverage
The latest PR already includes replay warnings in all three changelogs. The
implementation preserves result ordering, waits for all when_all children
before surfacing failure, and notifies a when_any parent only for the first
winner. No additional public API or minimum-Python-version change is needed.
The new tests exercise task objects directly. An executor-level regression
would additionally preserve the intended scheduling/replay boundary: successful
nested completion, replay of a history written by the corrected implementation,
and an unchanged cancellation-first history. The legacy-history mismatch above
is also useful to encode as an intentional compatibility test.
The automated review's empty-when_all([]) concern reproduces on both the merge
base and the PR. It is a pre-existing empty-input behavior, not a regression
introduced by completion propagation. Keep any empty-input semantic change
separate rather than silently expanding this fix.
WhenAllTaskandWhenAnyTaskmark themselves complete, but do not notify their parent composite.This means the following orchestration does not resume when all activities complete:
A current workaround is to wait for each leaf task with repeated
when_anycalls. The workaround preserves cancellation, but Durable Task Scheduler no longer shows the standard Fan-out group.Could composite completion propagate to the parent in the same way that
CompletableTask.complete()does?A regression test can compose
when_any([cancel, when_all([a, b])]), completeaandb, and verify that the outer task completes with the innerwhen_alltask.Compatibility and behavior-change considerations
PR #267 fixes the missing completion propagation unconditionally. The proposed
direction is to accept this intentional replay-compatibility break rather than
preserve the bug, provided the affected scenarios and upgrade guidance are
explicit. This is a behavioral/replay change, not a public signature change or
a new abstract-method requirement for third-party subclasses.
Affected histories
An existing instance can have a history like this:
when_any([cancel, when_all(tasks)])is pending.when_alldoes not notify the race.work, for example a
cleanupactivity.success branch instead. If that branch schedules
finishwhere historyrecorded
cleanup, replay fails withNonDeterminismError.Warning
A changed replay is not guaranteed to raise
NonDeterminismError. If bothbranches schedule the same activity name/ID but use different inputs or local
continuation state, the executor can accept the recorded action and continue
with the new branch. An old history using
record("cleanup")was reproducedcompleting with
"cleanup"on the merge base and"finish"with the fix.Do not rely on nondeterminism errors alone to identify affected instances.
This is not a blanket break for all orchestrations. Ordinary non-nested
composites retain their behavior. A cancellation-first history for the example
above still replays consistently. A history still blocked at the nested
composite, with no conflicting downstream continuation, can regain progress
when the orchestration next executes. Deploying the fix alone does not
necessarily wake an otherwise idle instance.
Provider and deployment implications
Both
durabletask.azuremanagedandazure-functions-durablev2 use the coreimplementation, including the Functions compatibility APIs
context.task_all()and
context.task_any(). Their current core constraints are open-ended:durabletask>=1.10.1anddurabletask[opentelemetry]>=1.10.1, respectively.Pinning only the provider therefore does not freeze this behavior; a fresh
dependency resolution can pick up the correction without a provider-version
change. A core major-version bump alone would not protect those existing
provider constraints.
durabletaskversion until the transition is deliberate.deployment/task hub and use a separate task hub for new instances.
the transition. Stuck instances may require deliberate recovery rather than
simply waiting to drain.
notes, including upgrades from earlier Functions v2 prereleases. Coordinate
package/dependency version changes in the release PR, not this fix.
PR scope and regression coverage
The latest PR already includes replay warnings in all three changelogs. The
implementation preserves result ordering, waits for all
when_allchildrenbefore surfacing failure, and notifies a
when_anyparent only for the firstwinner. No additional public API or minimum-Python-version change is needed.
The new tests exercise task objects directly. An executor-level regression
would additionally preserve the intended scheduling/replay boundary: successful
nested completion, replay of a history written by the corrected implementation,
and an unchanged cancellation-first history. The legacy-history mismatch above
is also useful to encode as an intentional compatibility test.
The automated review's empty-
when_all([])concern reproduces on both the mergebase and the PR. It is a pre-existing empty-input behavior, not a regression
introduced by completion propagation. Keep any empty-input semantic change
separate rather than silently expanding this fix.