| 1 |
Tue |
Sep 1 |
Introduction |
Zhiyao |
Lyon: Machine Hearing: An Emerging Field |
|
Course Project;
HW0 |
|
| 1 |
Thu |
Sep 3 |
Auditory Scene Analysis |
Zhiyao |
Bregman: ASA Book Chapter 1 |
Wang & Brown: CASA Book Chapter 1 |
HW1 |
|
| 2 |
Tue |
Sep 8 |
Signal Processing Review |
Zhiyao |
Mueller: Fundamentals of Music Processing Book Chapter 2 |
|
|
|
| 2 |
Thu |
Sep 10 |
Python Programming for Audio\
Slides |
Venkat |
librosa;
Audio Input Representations |
MIR with Python;
Python for Scientific Audio |
|
|
|
| 3 |
Tue |
Sep 15 |
Single Pitch Detection |
Zhiyao |
Cheveigne: CASA Book Chapter 2.1-2.3 |
Cheveigne & Kawahara: YIN;
Kim et al: CREPE;
Gfeller et al: SPICE;
Riou et al: PESTO |
HW2 |
HW1 |
| 3 |
Thu |
Sep 17 |
Human Auditory Sensation |
Zhiyao |
Yost: Hearing Book Chapter 11
Yost: Hearing Book Chapter 13 |
Patterson: Auditory Images;
Lyon et al: Sparse Auditory Representations;
Shamma: Encoding Sound Timbre in the Auditory System;
Wang & Shamma: Spectral Shape Analysis; |
|
|
| 4 |
Tue |
Sep 22 |
Rhythm Analysis |
Zhiyao |
Mueller: Fundamentals of Music Processing Book Chapter 6
Ellis: Beat Tracking by Dynamic Programming |
Klapuri et al: Meter Analysis
Heydari et al: BeatNet
Zhao et al: Beat Transformer
Foscarin et al: Beat This!
Desblancs et al.: Zero Note Mamba
Chang & Su: Beast |
HW3 |
HW2 |
| 4 |
Thu |
Sep 24 |
Timbre Representation |
Zhiyao |
Herrera-Boyer et al: Signal Processing Methods for Music Transcription Book Chapter 6;
Tzanetakis: Music Data Mining Book Chapter 2 |
Childers et al.: The Cepstrum;
Davis & Mermelstein: MFCC
Huang et al: Music Timbre Style Transfer
Hermansky & Morgan: RASTA
Wu et al: Transplayer |
|
|
| 5 |
Tue |
Sep 29 |
NMF Audio Modeling |
Zhiyao |
Smaragdis & Brown: NMF Polyphonic Music Transcription |
Lee & Seung: NMF |
|
|
| 5 |
Thu |
Oct 1 |
More on NMF |
Zhiyao |
Smaragdis et al.: PLCA |
Virtanen: Monaural Sound Source Separation |
HW4 |
HW3 |
| 6 |
Tue |
Oct 6 |
HMM Audio Modeling |
Zhiyao |
Rabiner: HMM |
Mysore: PhD Thesis Chapter 2 |
|
|
| 6 |
Thu |
Oct 8 |
Deep Learning for Audio
CIRC Intro; Bluehive Cheat Sheet |
Zhiyao |
Goodfellow et al.: Deep Learning Book Chapter 6 |
|
HW5 |
HW4 (due Sat) |
| 7 |
Tue |
Oct 13 |
NO CLASS: Fall Break |
|
How to write a paper?
How to give a talk?
How to make a poster? |
|
|
|
| 7 |
Thu |
Oct 15 |
Deep Learning Implementation
PyTorch 101
|
Venkat |
Goodfellow et al.: Deep Learning Book Chapter 9 |
Hinton et al.: DNN for Speech Recognition;
DNN for Speech Separation;
Huang et al: Singing Voice Separation by RNN | ;
|
Project Proposal |
| 8 |
Tue |
Oct 20 |
BP derivation |
Zhiyao |
Goodfellow et al.: Deep Learning Book Chapter 14 |
Schluter & Bock: Onset Detection by CNN;
Hamel & Eck: Music Feature Learning with DBN;
|
|
|
| 8 |
Thu |
Oct 22 |
Speech Technology |
Zhiyao |
Ravanelli et al: SpeechBrain, Park et al: Review of Speaker Diarization
ASVSpoof2019 |
Diarization Extended Reading
Kassis & Hengartner: Breaking Voice Authentication
Graves et al.: Connectionist Temporal Classification (CTC)
Yu et al.: Permutation Invariant Training (PIT) |
HW6 |
HW5 |
| 9 |
Tue |
Oct 27 |
Voice Conversion |
Zhiyao |
Sisman et al.: Overview; Qian et al.: AutoVC |
Sun et al: PPG; Li et al.: StarGANv2-VC; Yang et al: StreamVC |
|
|
| 9 |
Thu |
Oct 29 |
Multi-pitch Analysis |
Zhiyao |
Cheveigne: CASA Book Chapter 2 |
Klapuri: Harmonicity and Spectral Smoothness
Duan et al: Peak and Non-peak Region |
|
HW6 |
| 10 |
Tue |
Nov 3 |
Multi-pitch Analysis |
Zhiyao |
Duan et al: Multi-pitch Streaming |
Poliner & Ellis: Discriminative Model;
Sigtia et al.: Neural Network for Piano Transcription |
|
|
| 10 |
Thu |
Nov 5 |
Self Supervised Learning for Music Understanding |
Zhiyao |
Gfeller et al: SPICE;
Riou et al: PESTO;
Buisson et al.: SSL for Music Segmentation |
Balestriero et al.: Cookbook for SSL;
Gui et al.: Survey on SSL |
|
|
| 11 |
Tue |
Nov 10 |
TBD |
Students |
|
|
|
|
| 11 |
Thu |
Nov 12 |
TBD |
Students |
|
|
|
|
| 12 |
Tue |
Nov 17 |
TBD |
Students |
|
|
|
|
| 12 |
Thu |
Nov 19 |
Score-Informed Source Separation |
Zhiyao |
Dannenberg & Raphael: Alignment and Accompaniment;
Ewert et al: SISS Overview |
Ewert & Muller: Score-informed NMF;
Duan et al: Soundprism |
|
|
| 13 |
Tue |
Nov 24 |
Project Status Update in Zhiyao's Office |
Students |
Check Google Doc for Schedule |
|
|
Project Status Update |
| 13 |
Thu |
Nov 26 |
NO CLASS: Happy Thanksgiving!
| |
|
|
|
|
| 14 |
Tue |
Dec 1 |
Audio-Visual Scene Understanding |
Zhiyao |
Arandjelovic & Zisserman: Objects that Sound; Owens & Efros: AV Scene Analysis |
Arandjelovic & Zisserman: Look Listen and Learn; Zhao et al.: Sounds of Pixels |
|
|
| 14 |
Thu |
Dec 3 |
Multi-channel Source Localization and Separation |
Zhiyao |
Stern et al: CASA Book Chapter 5;
Yilmaz & Rickard: DUET |
Woodruff & Wang: Binaural Localization Reverberant Noisy |
|
|
| 15 |
Tue |
Dec 8 |
Interactive Music Systems |
Zhiyao |
Gifford et al.: Computational Systems for Music Improvisation |
Tatar & Pasquier: Music Agents |
|
|
| 15 |
Thu |
Dec 10 |
Music Generation |
TBD |
Benetatos et al.: BachDuet; Dhawiwal et al.: Jukebox |
Hadjeres et al.: DeepBach; Roberts et al.: Hierarchical Latent Vector Model; Jaques et al.: Generating Music with Reinforcement Learning; |
|
|
| 16 |
Sun |
Dec 20 |
Project Oral Presentations |
Students |
7:15PM-10PM in CSB 601 |
|
|
Project Report Final (due Mon);
Slides Final (due Mon) |