Introduction to Machine Learning - Lecture Slides

Lecture Title Topics Covered Reference PDF
Lec 1 Probability and Events Sample space; Events; Axioms of probability; Union bound; Inclusion-Exclusion principle; Conditional probability; Independence; Bayes rule. ProbabilityCourse.com - Chapters 1-2 View
Lec 2 Random Variables Discrete and continuous RVs; PMF, PDF, CDF; Joint and marginal; Independence; Correlation; Mean; Variance. ProbabilityCourse.com - Chapter 3 View
Lec 3 Random Variables (Cont.) Properties of variance; Laws of total expectation and variance; Entropy; Cross-entropy; Covariance matrix; Multivariate Gaussian. ProbabilityCourse.com - Chapters 4-5 View
Lec 4 Linear Algebra Norms; Inner products; Linear independence; Orthogonality; Span; Basis; Matrix multiplication as linear combination. Strang, Linear Algebra and Learning from Data (2019) - Chapter 1 View
Lec 5 Linear Algebra (Matrices) Linear maps; Column and null spaces; Rank-nullity; Orthogonal matrices; Eigenvalues and eigenvectors. Strang, Linear Algebra and Learning from Data (2019) - Chapter 1 View
Lec 6 Linear Algebra (EVD and SVD) Orthonormal eigen-basis; Coordinate transforms; Eigenvalue decomposition; Singular value decomposition. Strang, Linear Algebra and Learning from Data (2019) - Chapter 1 View
Lec 7 ML Paradigms and Metrics Supervised, unsupervised and reinforcement learning; MSE; MAE; R2; Accuracy; Precision; Recall; F-score. Murphy, Probabilistic Machine Learning: An Introduction - Section 5.1 View
Lec 8 Hypothesis Class and Bias-Variance Hypothesis class; Inductive bias; Generalization; Bias-variance decomposition; Bias-variance trade-off. ISL (An Introduction to Statistical Learning) - Chapter 2 View
Lec 9 Linear Regression One-hot and ordinal encoding; feature scaling (standardization, min-max); gradients and the chain rule; Least squares and its geometric interpretation (projection). Hastie et al., The Elements of Statistical Learning (ESL), Sec. 3.2 View
Lec 10 Regularized Least Squares Feature expansion; L2 Regularization (Ridge Regression); L1 Regularization (LASSO); Cross-validation. Hastie et al., The Elements of Statistical Learning (ESL book), Sec. 3.4.1 and 3.4.2 View
Lec 11 Logistic Regression Bayes Optimal Classifier; Sigmoid function; Maximum Likelihood; Cross-entropy; Gradient Descent. Hastie et al., The Elements of Statistical Learning (ESL book), Sec. 4.1 View
Lec 12 Softmax Regression and Support Vector Machines Softmax Regression; MAP Estimation; Gaussian Prior and L2 Regularization; Maximum-Margin Classification. Murphy, Probabilistic Machine Learning: An Introduction, Sec. 10.3 View
Lec 13 Support Vector Machines Soft-Margin SVM; Hinge Loss and Regularization Parameter C; SVM Dual; Kernel Trick; Murphy, Probabilistic Machine Learning: An Introduction, Sec. 17.3 View
Lec 14 Decision Trees Classification Trees; Regression Trees; Entropy; Gini Index; Murphy, Probabilistic Machine Learning: An Introduction, Sec. 18.1 View
Lec 15 Ensemble Methods: Bagging and Random Forests Cost-Complexity Pruning; Stacking; Committee of Experts; Bagging; Random Forests Murphy, Probabilistic Machine Learning: An Introduction, Sec. 18.2 to 18.4 View
Lec 16 Ensemble Methods: Boosting Techniques AdaBoost; Least Squares Boosting; Gradient Boosting Murphy, Probabilistic Machine Learning: An Introduction, Sec. 18.5 View
Lec 17 Dimensionality Reduction Prinicipal Component Analysis (PCA) Murphy, Probabilistic Machine Learning: An Introduction, Chapter 20 View
Lec 18 Clustering Techniques K-Means; Hierarchical Clustering Murphy, Probabilistic Machine Learning: An Introduction, Chapter 21 View
Lec 19 Gaussian Mixture Model MLE; GMM; EM Algorithm Murphy, Probabilistic Machine Learning: An Introduction, Chapter 21 View
Lec 20 Deep Neural Networks: Multi Layer Perceptron MLP Architecture; Activation Function; Back Propogation; Murphy, Probabilistic Machine Learning: An Introduction, Chapter 13 View
Lec 21 Practical Aspects of Training Deep Neural Networks SGD and its variants; Initialization; Batch Normalization Murphy, Probabilistic Machine Learning: An Introduction, Chapter 13 View
Lec 22 Convolutional Neural Network (CNN) Convolution; Pooling; CNN architecture; Murphy, Probabilistic Machine Learning: An Introduction, Chapter 14 View
Lec 23 Recurrent Neural Network (RNN) word embeddings; RNN; Bi-directional RNN; Encoder-Decoder Architecture Murphy, Probabilistic Machine Learning: An Introduction, Chapter 15 View
Lec 24 Attention Mechanism and Transformer Architecture Multi-Head Attention; Transformer Architecture Murphy, Probabilistic Machine Learning: An Introduction, Chapter 15 View