Time Line of the Lecture "Pattern Recognition"

  • Fundamentals
    • Speech recognition and understanding
    • Applications and system variants
    • Evaluation
  • Statistical speech recognition
    • Maximum a-posteriori (MAP) rule
    • Model simplification
    • Modeling
  • Conclusion and outlook

  • Motivation
  • Fundamentals
    • The „hidden“ part of the model
    • The inner family of random processes
  • Fundamental problems of Hidden Markov Models
    • Efficient calculation of sequence probabilities
    • Efficient calculation of the most probable sequence

  • Motivation
  • Basics of speaker verification und speaker identification
    • Preprocessing and segmentation
    • Codebook-based schemes
    • Schemes based on Gaussian mixture models
  • Model adaption
  • Discriminative approaches

  • Motivation
  • Fundamentals
    • Gaussian mixture models in practice
    • Generation of Gaussian mixture models
  • Applications in speech and audio processing
    • Bandwidth extension
    • Signal separation
    • Speaker recognition

  • Motivation
  • System concept
  • Extension of the excitation signal
    • Spectral shifting and modulation
    • Non-linear characteristics
  • Extension of the spectral envelope
    • Approaches using neural networks
    • Codebook-based approaches
    • Linear mapping
  • Examples

  • Motivation
  • Application examples
  • Cost function for the training of a codebook
  • LBG- and k-means algorithm
    • Basic schemes
    • Extensions
  • Combination with additional mapping schemes

  • Introduction
  • Features for speech and speaker recognition
    • Fundamental frequency
    • Spectral envelope
  • Representation of the spectral envelope
    • Predictor coefficients
    • Cepstral coefficients
    • Mel-filtered cepstral coefficients

  • Introduction
  • Characteristic of multi-microphone systems
  • Delay-and-sum structures
  • Filter-and-sum structures
  • Interference compensation
  • Audio examples and results
  • Outlook on postfilter structures

  • Generation and properties of speech signals
  • Wiener filter
  • Frequency-domain solution
  • Extensions of the gain rule
  • Extensions of the entire framework
  • Empirical mode decomposition

Recent Publications

T. O. Wisch, T. Kaak, A. Namenas, G. Schmidt: Spracherkennung in stark gestörten Unterwasserumgebungen, Proc. DAGA, Germany, 2018

S. Graf, T. Herbig, M. Buck, G. Schmidt: Low-Complexity Pitch Estimation Based on Phase Differences Between Low-Resolution Spectra, Proc. Interspeech, pp. 2316 -2320, 2017


