| |
Acoustics and Speech Technology
Speech is sound produced by the human vocal apparatus and shaped into the patterns of language. Air from the lungs passes the vocal folds, which vibrate to create voiced sounds, while the throat, mouth and nasal passages act as a resonant filter that emphasises certain frequencies. The resulting acoustic signal carries most of its useful energy across a limited band, which is why telephone systems have traditionally restricted the voice channel to a few kilohertz of bandwidth while still remaining intelligible.
To store, transmit or process speech electronically, the continuous acoustic waveform captured by a microphone must be converted into numbers. Pulse-code modulation (PCM) does this by sampling the signal at regular intervals and assigning each sample a numerical value, a process governed by the sampling rate and the number of bits per sample. Adaptive differential PCM (ADPCM) reduces the data needed by encoding the difference between successive samples and adapting its step size to the signal, achieving lower bit rates than plain PCM at comparable quality.
Speech coding goes further by exploiting the predictable structure of the voice. Because speech is produced by a known mechanism, coders can model the vocal tract and transmit a compact description rather than the raw waveform, which is how mobile telephones squeeze conversations into narrow channels. Throughout, engineers balance bandwidth, the signal-to-noise ratio (SNR) and loudness so that the reconstructed voice remains clear. Pitch, the rate at which the vocal folds vibrate, conveys intonation and helps distinguish one speaker from another.
Frequently asked questions
- What is the difference between PCM and ADPCM?
- PCM stores the full value of each sample, while ADPCM encodes only the change from one sample to the next and adapts its step size, which lowers the data rate for similar quality.
- Why is telephone speech limited to a narrow band?
- Most of the energy needed to understand speech lies in a few kilohertz, so restricting the channel saves bandwidth while keeping the voice intelligible.
- How can speech be compressed so much for mobile phones?
- Speech coders model the way the voice is produced and transmit a compact description of that model rather than the raw waveform, allowing very low bit rates.
|
|
Speech technology
related topics:
Audio file formats,
Digital audio,
Digital audio samples,
Speech -anatomy,
Telephony,
VoIP - internet |
|
Audio compression formats
A comparison of Internet audio compression formats, Audio format 16-bit PCM
G.711 mu-law 32Kbps MPEG-1 IMA/DVI ADPCM GSM 06.10 InterWave VSC112 TrueSpeech
8.5 RealAudio v1.0 ToolVox for the Web |
|
Human language
technology Survey of the State of the Art in Human Language Technology |
|
Phonetics and speech
processing phonology, Phonetics and speech processing, music and sound waves |
|
Phonetics topics |
|
Sopranos: resonance tuning and vowel changes |
| Speech
acoustics speech acoustics |
|
Speech and audio coding Wideband Speech and Audio Coding, ISO/MPEG
Standardization, Operating Modes, Frequency Mapping, Quantization and Bit
Allocation, Pre-Echo Control, ISO/MPEG Layers |
|
Speech
coding principles of speech coding, narrowband speech codecs. Such codecs
are used to give an efficient digital representation of telephone bandwidth
speech. Often the speech is bandlimited to between 200 and 3400 Hz, and is
sampled at 8 kHz |
|
Speech
coding PCM speech, DPCM, ADPCM, quantization, Variable length Coding, VLC,
Huffman coding, Predictive Coding (DPCM), delta modulation,
pdf file |
|
Speech
coding speech coding,
pdf file |
|
Speech
Processing Speech recognition
speech processing in time and frequency domain, Markov models |
|
The MARY Text-to-Speech System
MARY is a Text-to-Speech Synthesis System for German, English and Tibetan,
a tip |
|
Vocoders RPE-LTP: Regular Pulse Excited-Long Term Prediction, GSM, Digital
speech model, ppt file |
|
Voice coder demonstration
Audio and Voice Coder Demonstration, Code-Excited Sample-by-Sample Gain Adaptive
Trellis Coding, Sample-by-Sample Adaptive Differential Vector Quantization,
Harmonic Excited Linear Prediction |
| Voice coding
algorithms (click left on tutorials, scroll down) PCM, ADPCM, Adaptive Differential Pulse Code Modulation, loudness |
| VOICE DIGITIZATION AND VOICE/DATA INTEGRATION
pdf file |
|
|
Home
|
Site Map
|
Email: support[at]karadimov.info
Last updated on:
2026-06-24
|
Copyright © 2011-2021 Educypedia.
https://educypedia.org
|
| |
|
|