| Lec 1 |
Probability and Events |
Sample space; Events; Axioms of probability; Union bound; Inclusion-Exclusion principle; Conditional probability;
Independence; Bayes rule.
|
ProbabilityCourse.com - Chapters 1-2 |
View |
| Lec 2 |
Random Variables |
Discrete and continuous RVs; PMF, PDF, CDF; Joint and marginal;
Independence; Correlation; Mean; Variance.
|
ProbabilityCourse.com - Chapter 3 |
View |
| Lec 3 |
Random Variables (Cont.) |
Properties of variance; Laws of total expectation and variance;
Entropy; Cross-entropy; Covariance matrix; Multivariate Gaussian.
|
ProbabilityCourse.com - Chapters 4-5 |
View |
| Lec 4 |
Linear Algebra |
Norms; Inner products; Linear independence; Orthogonality;
Span; Basis; Matrix multiplication as linear combination.
|
Strang, Linear Algebra and Learning from Data (2019) - Chapter 1
|
View |
| Lec 5 |
Linear Algebra (Matrices) |
Linear maps; Column and null spaces; Rank-nullity;
Orthogonal matrices; Eigenvalues and eigenvectors.
|
Strang, Linear Algebra and Learning from Data (2019) - Chapter 1
|
View |
| Lec 6 |
Linear Algebra (EVD and SVD) |
Orthonormal eigen-basis; Coordinate transforms; Eigenvalue decomposition;
Singular value decomposition.
|
Strang, Linear Algebra and Learning from Data (2019) - Chapter 1
|
View |
| Lec 7 |
ML Paradigms and Metrics |
Supervised, unsupervised and reinforcement learning;
MSE; MAE; R2; Accuracy; Precision; Recall; F-score.
|
Murphy, Probabilistic Machine Learning: An Introduction - Section 5.1
|
View |
| Lec 8 |
Hypothesis Class and Bias-Variance |
Hypothesis class; Inductive bias; Generalization;
Bias-variance decomposition; Bias-variance trade-off.
|
ISL (An Introduction to Statistical Learning) - Chapter 2
|
View |
| Lec 9 |
Linear Regression |
One-hot and ordinal encoding; feature scaling (standardization, min-max);
gradients and the chain rule; Least squares and its geometric interpretation (projection).
|
Hastie et al., The Elements of Statistical Learning (ESL), Sec. 3.2
|
View |
| Lec 10 |
Regularized Least Squares |
Feature expansion; L2 Regularization (Ridge Regression);
L1 Regularization (LASSO); Cross-validation.
|
Hastie et al., The Elements of Statistical Learning (ESL book), Sec. 3.4.1 and 3.4.2
|
View |
| Lec 11 |
Logistic Regression |
Bayes Optimal Classifier; Sigmoid function; Maximum Likelihood;
Cross-entropy; Gradient Descent.
|
Hastie et al., The Elements of Statistical Learning (ESL book), Sec. 4.1
|
View |
| Lec 12 |
Softmax Regression and Support Vector Machines |
Softmax Regression;
MAP Estimation;
Gaussian Prior and L2 Regularization;
Maximum-Margin Classification.
|
Murphy, Probabilistic Machine Learning: An Introduction, Sec. 10.3
|
View |
| Lec 13 |
Support Vector Machines
|
Soft-Margin SVM;
Hinge Loss and Regularization Parameter C; SVM Dual;
Kernel Trick;
|
Murphy, Probabilistic Machine Learning: An Introduction, Sec. 17.3
|
View
|
| Lec 14 |
Decision Trees
|
Classification Trees;
Regression Trees; Entropy;
Gini Index;
|
Murphy, Probabilistic Machine Learning: An Introduction, Sec. 18.1
|
View
|
| Lec 15 |
Ensemble Methods: Bagging and Random Forests
|
Cost-Complexity Pruning;
Stacking; Committee of Experts;
Bagging; Random Forests
|
Murphy, Probabilistic Machine Learning: An Introduction, Sec. 18.2 to 18.4
|
View
|
| Lec 16 |
Ensemble Methods: Boosting Techniques
|
AdaBoost;
Least Squares Boosting; Gradient Boosting
|
Murphy, Probabilistic Machine Learning: An Introduction, Sec. 18.5
|
View
|
| Lec 17 |
Dimensionality Reduction
|
Prinicipal Component Analysis (PCA)
|
Murphy, Probabilistic Machine Learning: An Introduction, Chapter 20
|
View
|
| Lec 18 |
Clustering Techniques
|
K-Means; Hierarchical Clustering
|
Murphy, Probabilistic Machine Learning: An Introduction, Chapter 21
|
View
|
| Lec 19 |
Gaussian Mixture Model
|
MLE; GMM; EM Algorithm
|
Murphy, Probabilistic Machine Learning: An Introduction, Chapter 21
|
View
|
| Lec 20 |
Deep Neural Networks: Multi Layer Perceptron
|
MLP Architecture; Activation Function; Back Propogation;
|
Murphy, Probabilistic Machine Learning: An Introduction, Chapter 13
|
View
|
| Lec 21 |
Practical Aspects of Training Deep Neural Networks
|
SGD and its variants; Initialization; Batch Normalization
|
Murphy, Probabilistic Machine Learning: An Introduction, Chapter 13
|
View
|
| Lec 22 |
Convolutional Neural Network (CNN)
|
Convolution; Pooling; CNN architecture;
|
Murphy, Probabilistic Machine Learning: An Introduction, Chapter 14
|
View
|
| Lec 23 |
Recurrent Neural Network (RNN)
|
word embeddings; RNN; Bi-directional RNN; Encoder-Decoder Architecture
|
Murphy, Probabilistic Machine Learning: An Introduction, Chapter 15
|
View
|
| Lec 24 |
Attention Mechanism and Transformer Architecture
|
Multi-Head Attention; Transformer Architecture
|
Murphy, Probabilistic Machine Learning: An Introduction, Chapter 15
|
View
|