Skip to content

add streaming beam decoding (malsd) to NeMo inference buffered rnnt - #15827

Open
lilithgrigoryan wants to merge 10 commits into
mainfrom
lgrigoryan/streaming-beam-search-niva-buffered-rnnt
Open

lilithgrigoryan wants to merge 10 commits into
mainfrom
lgrigoryan/streaming-beam-search-niva-buffered-rnnt

Conversation

@lilithgrigoryan

@lilithgrigoryan lilithgrigoryan commented Jun 24, 2026 •

Copy link
Copy Markdown
Collaborator

Important

The Update branch button must only be pressed in very rare occassions.
An outdated branch is never blocking the merge of a PR.
Please reach out to the automation team before pressing that button.

What does this PR do ?

Adds MALSD batched beam search decoding to the buffered RNNT streaming inference pipeline, including per-token confidence support.

Collection: ASR

Changelog

  • Support MALSD batched beam search in the buffered RNNT streaming pipeline, selected with asr.decoding.strategy=malsd_batch. Greedy decoding is unchanged.
  • Carry the full beam across chunks and collapse it to a single hypothesis only at end of utterance, using the same length-normalized ranking as offline decoding.
  • Support per-token confidence for streaming beam search.
  • Fix confidence in batched MALSD beam search (RNNT and TDT) being computed from LM-fused scores instead of acoustic log-probabilities, which could push values outside [0, 1].
  • Add an asr.decoding.beam section to the buffered RNNT config, with n-gram LM and phrase boosting disabled by default.

Usage

  • Beam search is selected with asr.decoding.strategy=malsd_batch. The asr.decoding.beam.* options mirror the offline MALSD decoding config. Setting asr.decoding.beam.preserve_frame_confidence=true adds per-token confidence to the output.
python examples/asr/asr_streaming_inference/asr_streaming_infer.py \
  --config-path=examples/asr/conf/asr_streaming_inference \
  --config-name=buffered_rnnt.yaml \
  asr.model_name=nvidia/parakeet-rnnt-1.1b \
  audio_file=<path to manifest>.json \
  output_filename=<path to output>.jsonl \
  calculate_wer=true \
  asr.decoding.strategy=malsd_batch \
  asr.decoding.beam.beam_size=4 \
  asr.decoding.beam.allow_cuda_graphs=true \
  asr.decoding.beam.preserve_frame_confidence=true \
  asr.decoding.beam.ngram_lm_model=<path to NGPU-LM>.nemo \
  asr.decoding.beam.ngram_lm_alpha=0.3 \
  streaming.chunk_size=1.12 \
  streaming.left_padding_size=5.6 \
  streaming.right_padding_size=1.12 \
  streaming.batch_size=256

GitHub Actions CI

The Jenkins CI system has been replaced by GitHub Actions self-hosted runners.

The GitHub Actions CI will run automatically when the "Run CICD" label is added to the PR.
To re-run CI remove and add the label again.
To run CI on an untrusted fork, a NeMo user with write access must first click "Approve and run".

Before your PR is "Ready for review"

Pre checks:

  • Make sure you read and followed Contributor guidelines
  • Did you write any new necessary tests?
  • Did you add or update any necessary documentation?
  • Does the PR affect components that are optional to install? (Ex: Numba, Pynini, Apex etc)
    • Reviewer: Does the PR have correct import guards for all optional libraries?

PR Type:

  • New Feature
  • Bugfix
  • Documentation

If you haven't finished some of the above items you can still open "Draft" PR.

Who can review?

Anyone in the community is free to review the PR once the checks have passed.
Contributor guidelines contains specific people who can review PRs to various areas.

Additional Information

  • Related to # (issue)

Signed-off-by: lilithgrigoryan <lgrigoryan@nvidia.com>
Signed-off-by: lilithgrigoryan <lgrigoryan@nvidia.com>
…/streaming-beam-search-niva-buffered-rnnt
Comment thread nemo/collections/asr/inference/pipelines/buffered_rnnt_pipeline.py Fixed
Signed-off-by: lilithgrigoryan <lgrigoryan@nvidia.com>
Signed-off-by: lilithgrigoryan <lgrigoryan@nvidia.com>
Signed-off-by: lilithgrigoryan <lgrigoryan@nvidia.com>
Signed-off-by: lilithgrigoryan <lgrigoryan@nvidia.com>
Signed-off-by: lilithgrigoryan <lgrigoryan@nvidia.com>
Comment thread nemo/collections/asr/parts/utils/rnnt_utils.py Fixed
…/streaming-beam-search-niva-buffered-rnnt
Signed-off-by: lilithgrigoryan <lgrigoryan@nvidia.com>
@github-actions

Copy link
Copy Markdown
Contributor

[🤖]: Hi @lilithgrigoryan 👋,

We wanted to let you know that a CICD pipeline for this PR just finished successfully.

So it might be time to merge this PR or get some approvals.

@naymaraq naymaraq left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I’ve added several small comments. Please take a look.
Also, if possible, please add the following ablation results to the PR:

  • Greedy inference before vs. after your changes
  • Greedy vs. beam search vs. beam search with LM fusion using your branch
  • Beam search consistency when changing the EoU threshold, including disabling it

"""Register per-stream biasing models and return decode-time model ids."""
if self.decoding_computer is None or not self.decoding_computer.per_stream_biasing_enabled:
if any(state.has_biasing_request() for state in states):
logging.warning(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Question:
should we rise an error here?

curr_state.timestamp_offset += self.tokens_per_frame_float
ready_state_ids.update(ready_states)

if self.beam_decoder_computer is not None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we join two loops into one?

# See the License for the specific language governing permissions and
# limitations under the License.

from __future__ import annotations

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we don’t need annotations here as long as the StreamingState and Hypothesis imports work correctly.

"""Refresh publish tokens; on EOU fold utterance into ``cumulative_*`` and clear ``partial_*``."""
cum_tokens, cum_ts, cum_conf = self._get_tokens()
if cum_tokens:
start = max(0, min(int(self._cumulative_tokens_len), len(cum_tokens)))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is the int() necessary here?

This branch was previously deployed

2 inactive deployments
public — 0392df8a Deployed Aug 27, 2026 by copy-pr-bot[bot] via release / finalize / notify #1242
test — 0392df8a Deployed Aug 27, 2026 by copy-pr-bot[bot] via cicd-wait-in-queue #20816
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants