IIT Madras Speech Technology Lectures
BSEE4001122 lectures12 weeks
Weeks
- Week 16 lecturesCourse introduction · Digital signal processing fundamentals | sampling, quantization & fourier basics · Fourier series | speech production & phoneme representation · Speech production | fourier, z-transform & digital filters · Revision | speech perception, masking & cepstral analysis · Waveforms | source-filter model, spectrograms & human hearing
- Week 26 lecturesSpeech production, perception & frequency analysis · Perceptual masking, cepstrum & filtering | speech analysis · Feature extraction for speech processing | part 01 · MFCC | liftering, mel scale & feature extraction for speech · MFCC · Gaussian review | binary classification & multivariate distributions
- Week 37 lecturesGaussian model for binary classification problem · Error analysis of gaussian model | mel filter banks · Introduction to bi variate gaussian model | gaussian mixture models · Mixture of gaussians introduction · Mixture of gaussians | binary classification & error trade-offs · Parameter estimation of GMM · Vector quantization introduction
- Week 46 lecturesVector quantization | data representation, clustering & finite precision · Markov model | forecasting, greedy algorithm & viterbi · Markov chain – example | sequential modeling & weather prediction · Predicting weather sequence | sequence modeling & real-world applications · Viterbi algorithm introduction · Viterbi algorithm | intuitive explanation
- Week 55 lecturesHidden markov model | weather example, transition matrices & viterbi motivation · Review of the weather prediction example · Efficiency of viterbi algorithm | hidden markov models, observations & weather example · Review of HMM weather & grass example · Forward algorithm for HMM
- Week 68 lecturesReview of speech feature extraction · Review of a pattern classification problem · Limited vocabulary speech recognition · Review of limited vocabulary speech recognition · Introduction to HMM GMM model for speech recognition | part 1 · Introduction to HMM GMM model for speech recognition | part 2 · HMM applied to speech | limited vocabulary recognition with yes/no acoustic models · HMM applied to speech | yes/no word recognition with HMM-GMM acoustic models
- Week 77 lecturesDifferent types of HMM models | three-state phoneme, triphone & state tying · Lexicon / pronunciation dictionary · Introduction to language model | artificial neurons & perceptrons in neural networks · N gram language model | history & evolution of neural networks in speech processing · Introduction to deep neural network · Deep learning basics | part 1 | n-gram language models & perplexity · Deep learning basics | part 2 | activation functions & non-linearities in neural networks
- Week 84 lecturesFeed forward neural network | part 1 | backpropagation, softmax & gradient descent · Revision of feed forward neural networks · Introduction to CTC | part 1 · Introduction to CTC | part 2 | neural networks, softmax & automated classification
- Week 95 lecturesCTC recap · Recurrent neural networks · Application in language modelling · Disadvantages with RNN · Disadvantages with RNN | language models, one-hot encoding
- Week 106 lecturesIntroduction to word2vec | word embeddings, nlp & beyond · Sequence to sequence modelling | ctc, encoder-decoder & attention · Introduction to CNN | signal processing foundations for convolutional neural networks · CNN example | convolution operations explained (1d, 2d, 3d) · CNN for speech · Sequence to sequence networks | CNNs for digit recognition & feature extraction
- Week 116 lecturesTransformer introduction | self-attention, queries, keys & values explained · Transformer – self attention | encoder-decoder, multi-head attention & training trade-offs · Transformer – training | self-supervised learning & autoencoders for speech · Transformers for speech · Introduction to self supervised learning · BERT
- Week 127 lecturesIntroduction to text to speech synthesis · End to end speech synthesis: auto regressive networks · End to end speech synthesis: non auto regressive networks · Introduction to speaker verification · Speaker classification architectures for speaker verification · Speaker verification scoring techniques & evaluation metrics · Speaker diarization