Awesome Testing

Models · reviewed · reviewed Oct 6, 2026 · 4 min

How does AI recognize speech?

Speech recognition encodes sampled audio and uses a learned model and decoder to produce text. Recording quality, missing segments, language, and decoding affect the result. A fluent transcript is a prediction, not proof that every word was heard correctly.

The input is a signal, not a sentence

A microphone records changes in air pressure as a sequence of numerical samples. A sampling rate describes how often a value is recorded. It does not count words, phonemes, or text tokens.

Our teaching recording uses 16,000 samples per second and one audio channel. The authored script is “Please send the report tomorrow.” The voice is synthetic, generated locally with macOS Samantha; no real person's recording or remote transcription service is involved.

The waveform carries timing, pitch, pauses, noise, and overlapping acoustic information. It does not arrive with spaces between words. A recognizer must learn how patterns in that signal relate to the text it produces.

Change what the recognizer would receive

Listen to the clean clip, then select its noisy, cropped, or silent variant. The peak envelope shows amplitude across the selected file. The frequency view samples short windows and displays how energy is distributed across selected frequencies over time.

These figures are computed from the exact WAV files supplied to the player. No recognizer runs in the experiment, and the known script remains separate from any predicted transcript. Predict which information disappears when only the first 0.75 seconds remain.

Explore the mechanism

Change the recording, inspect the signal

Known reference script: Please send the report tomorrow. This is the recording’s script, not an inferred transcript of the selected variant.

Peak amplitude across the selected clip
Sparse frequency view: time →, low to high frequency ↑
Available duration
1.75 s
PCM samples
27974
RMS amplitude
0.170

The clean waveform contains the authored utterance; a learned decoder would still be needed to recognize it.

Compare a separate toy CTC alignment

This invented letter path is unrelated to the recording. First collapse adjacent repeats, then remove blanks (_).

b b _ o o _ o _ k k → book.

b b _ o o o _ k k → bok. A blank between the two o groups preserves two letters. This is a CTC path transformation, not Whisper’s decoding algorithm.

The authored reference script was synthesized locally with macOS Samantha. The four WAV files and their computed figures are fixed assets. Noise uses seed 7; the cropped clip retains only 0.75 seconds. No speech recognizer runs. The sparse linear-frequency plot is a teaching projection, not Whisper’s Mel frontend or a recognition-confidence map.

The crop contains fewer samples and less elapsed audio. The fact that we know the original script does not establish that its later words are present in the cropped input. Added noise changes the signal without revealing how a particular model would transcribe it. Silence supplies no spoken evidence.

The frequency plot is an illustrative linear-frequency projection. It is not a word map, a confidence score, or the complete frontend of a deployed recognizer. A region becoming darker means more energy under the display rule; it does not mean a word has been recognized.

Features prepare the input for a learned model

A spectrogram examines frequency content in successive short windows. Different frontends use different representations. A Mel spectrogram groups frequency information using a perceptually motivated scale; a learned waveform encoder can instead discover useful representations from audio samples.

Whisper provides one concrete architecture: it processes a log-Mel representation with an encoder and generates text with an autoregressive decoder. Its task format can distinguish transcription from translation. These operations are different: writing the spoken language and translating it into another language do not have the same target.

wav2vec 2.0 illustrates another route. It learns contextual speech representations from waveform input, then adapts the model for recognition with labelled examples. Neither approach turns one audio frame directly into one word; useful representations depend on neighbouring information and training.

A decoder must handle timing and alternatives

One utterance may occupy many acoustic time steps, while its transcript contains relatively few symbols. Connectionist Temporal Classification, or CTC, provides one way to learn across possible alignments without requiring an exact timestamp for every training label.

In the separate toy path shown by the experiment, adjacent repeated symbols collapse first, then blanks disappear. Two o groups separated by a blank can preserve the two letters in “book”; one uninterrupted o group cannot. This is an invented CTC example, unrelated to the recorded sentence.

Whisper’s autoregressive decoder uses a different procedure: generate a text token, incorporate that prefix, and continue conditioned on the encoded audio. A language-like continuation can be plausible even when the audio evidence is weak. Text fluency therefore remains separate from transcription fidelity.

The transcript is an intermediate result

Suppose the recording contains a negation or a product code that determines the requested action. A transcript that drops “not” or changes one digit can remain grammatically natural while reversing the meaning. Preserve the original audio reference and inspect consequential uncertainty instead of treating parseable text as a complete record.

Long recordings add segmentation and continuity decisions. Cutting through speech, losing a chunk, or mixing speakers can remove information before decoding. Speaker identification, timestamps, text normalization, and downstream instructions also have distinct owners; a speech-recognition result does not automatically solve them all.

Recognition maps audio evidence to predicted text. Speech synthesis maps a specification such as text toward generated audio. The synthetic voice used to create this fixture performs the latter operation; the published experiment does not secretly perform the former.

Recording and figure provenance

The reference was produced with the installed Samantha English voice at rate 150, then converted to 16 kHz mono 16-bit PCM. Its exact bytes are versioned. The noisy file adds a fixed uniform draw with seed 7 and amplitude 0.12, bounded to the PCM range; the crop keeps 0.75 seconds; silence keeps the original duration with zero samples.

The figures use a 120-bin peak envelope and 48 sampled, 400-sample Hann-windowed frequency analyses. Thirty-two selected frequency bins extend to 8 kHz. A fixed logarithmic display range maps power to color consistently across variants; it is not a Mel filter bank or a learned encoder.

The versioned fixture generator records duration, sample count, RMS amplitude, feature arrays, and WAV hashes. Descriptions identify the variants; the cropped description deliberately does not supply the full script as a transcript. Playback is manual and stops when the selected audio element is replaced.

Sources and further reading

  1. 01
    Robust Speech Recognition via Large-Scale Weak SupervisionRadford et al. · research · published Dec 6, 2022 · source checked Oct 6, 2026

    A concrete speech-recognition architecture uses log-Mel audio features, an encoder, and a task-conditioned autoregressive text decoder.

  2. 02
    wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsBaevski et al. · research · published Jun 20, 2020 · source checked Oct 6, 2026

    Learned waveform representations and downstream CTC recognition distinguish acoustic features, pretraining, adaptation, and text decoding.

  3. 03
    Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural NetworksGraves et al. · research · source checked Oct 6, 2026

    Original alignment-path formulation with repeated labels and blanks supports the separate disclosed CTC letter example.