Benchmarking#
process() is a blocking call that takes a complete audio clip and returns
the final transcript. It is suited for benchmarking throughput and accuracy against a fixed audio
dataset.
For real-time or streaming use, see Overview.
Asr.process()#
Asr.process(data: bytes) -> AsrTranscript
Pushes all of data through the ASR pipeline, waits for the neural network to finish, and returns
the transcript. The call blocks until the entire input has been processed; it does not return
partial results.
process() can be called multiple times on the same
Asr instance. Each call is self-contained: no audio state carries over
between clips, so you can process a batch of clips back to back.
Parameters#
Parameter |
Type |
Description |
|---|---|---|
|
|
PCM audio encoded as a little-endian 16-bit byte array at 16,000 Hz. See Input format for format requirements. |
Return value#
Returns an AsrTranscript. Call .text to get the final string.
Example#
from pathlib import Path
from abr_sdk.asr import Asr
LIBRARY_PATH = "/path/to/niagara-38m-live.en/libniagara_38m_live.so"
CLIP_PATHS = ["clip1.pcm", "clip2.pcm"]
with Asr(LIBRARY_PATH) as asr:
for path in CLIP_PATHS:
transcript = asr.process(Path(path).read_bytes())
print(transcript.text)
Each .pcm file must contain raw PCM audio in the format the Asr instance expects. See
Input format for details on sample rate, bit depth, and channel layout.
Errors#
Exception |
When raised |
|---|---|
|
Called on a closed |
|
The underlying C library returned a failure status. Check the log output for details. |