Educypediathe educational encyclopedia
Electronics-theory
Analog
Audio - acoustics
Audio - electronics
Audio - loudspeakers
Components-active
Components-passive
Component - sensors
Digital - electronics
Digital - I2C - I2S
Digital - Programming
Electricity High Voltage
Electricity - machines
Electricity - theory
General overview
Miscellaneous
Optics
Power control systems
Power electronics
RF - antennas
RF - antenna - WLAN
RF - communication
RF - radio - tuning
Telephony
Tubes
TV-video-DVD
 
Utilities - tools
Animations & applets
Application notes
Cabling
Calculators
Circuits
Databank - tables
Datasheets
Measurement
Repair
Software - electronics
 
Local sitemap
Sitemap

  
 

Acoustics and Speech Technology

Speech is sound produced by the human vocal apparatus and shaped into the patterns of language. Air from the lungs passes the vocal folds, which vibrate to create voiced sounds, while the throat, mouth and nasal passages act as a resonant filter that emphasises certain frequencies. The resulting acoustic signal carries most of its useful energy across a limited band, which is why telephone systems have traditionally restricted the voice channel to a few kilohertz of bandwidth while still remaining intelligible.

To store, transmit or process speech electronically, the continuous acoustic waveform captured by a microphone must be converted into numbers. Pulse-code modulation (PCM) does this by sampling the signal at regular intervals and assigning each sample a numerical value, a process governed by the sampling rate and the number of bits per sample. Adaptive differential PCM (ADPCM) reduces the data needed by encoding the difference between successive samples and adapting its step size to the signal, achieving lower bit rates than plain PCM at comparable quality.

Speech coding goes further by exploiting the predictable structure of the voice. Because speech is produced by a known mechanism, coders can model the vocal tract and transmit a compact description rather than the raw waveform, which is how mobile telephones squeeze conversations into narrow channels. Throughout, engineers balance bandwidth, the signal-to-noise ratio (SNR) and loudness so that the reconstructed voice remains clear. Pitch, the rate at which the vocal folds vibrate, conveys intonation and helps distinguish one speaker from another.

Frequently asked questions

What is the difference between PCM and ADPCM?
PCM stores the full value of each sample, while ADPCM encodes only the change from one sample to the next and adapts its step size, which lowers the data rate for similar quality.
Why is telephone speech limited to a narrow band?
Most of the energy needed to understand speech lies in a few kilohertz, so restricting the channel saves bandwidth while keeping the voice intelligible.
How can speech be compressed so much for mobile phones?
Speech coders model the way the voice is produced and transmit a compact description of that model rather than the raw waveform, allowing very low bit rates.





Speech technology  related topics: Audio file formats, Digital audio, Digital audio samples, Speech -anatomy, Telephony, VoIP - internet
Audio compression formats A comparison of Internet audio compression formats, Audio format 16-bit PCM G.711 mu-law 32Kbps MPEG-1 IMA/DVI ADPCM GSM 06.10 InterWave VSC112 TrueSpeech 8.5 RealAudio v1.0 ToolVox for the Web
Human language technology Survey of the State of the Art in Human Language Technology
Phonetics and speech processing phonology, Phonetics and speech processing, music and sound waves
Phonetics topics
Sopranos: resonance tuning and vowel changes
Speech acoustics speech acoustics
Speech and audio coding Wideband Speech and Audio Coding, ISO/MPEG Standardization, Operating Modes, Frequency Mapping, Quantization and Bit Allocation, Pre-Echo Control, ISO/MPEG Layers
Speech coding principles of speech coding, narrowband speech codecs. Such codecs are used to give an efficient digital representation of telephone bandwidth speech. Often the speech is bandlimited to between 200 and 3400 Hz, and is sampled at 8 kHz
Speech coding PCM speech, DPCM, ADPCM, quantization, Variable length Coding, VLC, Huffman coding, Predictive Coding (DPCM), delta modulation, pdf file
Speech coding speech coding, pdf file
Speech Processing Speech recognition speech processing in time and frequency domain, Markov models
The MARY Text-to-Speech System MARY is a Text-to-Speech Synthesis System for German, English and Tibetan, a tip
Vocoders RPE-LTP: Regular Pulse Excited-Long Term Prediction, GSM, Digital speech model, ppt file
Voice coder demonstration Audio and Voice Coder Demonstration, Code-Excited Sample-by-Sample Gain Adaptive Trellis Coding, Sample-by-Sample Adaptive Differential Vector Quantization, Harmonic Excited Linear Prediction
Voice coding algorithms (click left on tutorials, scroll down) PCM, ADPCM, Adaptive Differential Pulse Code Modulation, loudness
VOICE DIGITIZATION AND VOICE/DATA INTEGRATION pdf file

Home | Site Map | Email: support[at]karadimov.info

Last updated on: 2026-06-24 | Copyright © 2011-2021 Educypedia.

https://educypedia.org

 

 

 

 
Powered by ITCom Solutions