|
|
MSDM 5054: Statistical Machine Learning
Fall 2026
|
Synopsis
This course covers several topics in statistical machine learning:
- 1. supervised learning (linear and nonlinear models, e.g. trees, support vector machines, deep neural networks),
- 2. unsupervised learning (dimensionality reduction, cluster trees, generative models),
- 3. reinforcement learning (markov decision process, deep rl).
Prerequisite: Some preliminary course on (statistical) machine learning, applied statistics, and deep learning will be helpful.
Instructors:
Yuan Yao
Time and Place:
Tu 6:30-9:20pm, Lecture Theater D (LTD), HKUST
Reference (参考教材)
An Introduction to Statistical Learning, with applications in R (ISLR). By James, Witten, Hastie, and Tibshirani
ISLR-python, By Jordi Warmenhoven.
ISLR-Python: Labs and Applied, by Matt Caudill.
Manning: Deep Learning with Python, by Francois Chollet [GitHub source in Python 3.6 and Keras 2.0.8]
MIT: Deep Learning, by Ian Goodfellow, Yoshua Bengio, and Aaron Courville
Tutorials: preparation for beginners
Python-Numpy Tutorials by Justin Johnson
scikit-learn Tutorials: An Introduction of Machine Learning in Python
Jupyter Notebook Tutorials
PyTorch Tutorials
Deep Learning: Do-it-yourself with PyTorch, A course at ENS
Tensorflow Tutorials
MXNet Tutorials
Theano Tutorials
The Elements of Statistical Learning (ESL). 2nd Ed. By Hastie, Tibshirani, and Friedman
statlearning-notebooks, by Sujit Pal, Python implementations of the R labs for the StatLearning: Statistical Learning online course from Stanford taught by Profs Trevor Hastie and Rob Tibshirani.
Homework and Projects:
TBA (To Be Announced)
Teaching Assistant:
Email: Ms. WANG, Ruihe < ruihe.wang (add "AT connect DOT ust DOT hk" afterwards) >
Schedule
| Date |
Topic |
Instructor |
Scriber |
| 09/07/2026, Mon |
Lecture 01: A Historic Overview of AI and Statistical Machine Learning. [ slides (pdf) ]
|
Y.Y. |
|
| 09/14/2026, Mon |
Lecture 02: Supervised Learning: linear regression and classification [ slides ]
|
Y.Y. |
|
| 09/21/2026, Mon |
Lecture 03: Model Assessment and Selection: Subset, Ridge, Lasso, and PCR [ slides ] and Project 1 [ pdf ]
[Reference]:
- To view .ipynb files below, you may try [ Jupyter NBViewer]
- Python Notebook for Model Selection (Subset, Ridge, Lasso, and Principal Component Regression)
[ Selection.ipynb ]
- Kaggle: Home Credit Default Risk [ link ]
- Kaggle: M5 Forecasting - Accuracy, Estimate the unit sales of Walmart retail goods.
[ link ]
|
Y.Y. |
|
| 09/28/2026, Mon |
Lecture 04: Moving beyond Linearity [ slides ]
[ Seminar ]
- Title: A Statistical View on Implicit Regularization: Gradient Descent Dominates Ridge [ slides ][ link ]
- Speaker: Dr. Jingfeng WU, UC Berkeley
- Abstract: A key puzzle in deep learning is how simple gradient methods find generalizable solutions without explicit regularization. This talk discusses the implicit regularization of gradient descent (GD) through the lens of statistical dominance. Using linear regression as a clean proxy, we present three surprising findings.
First, GD dominates ridge regression: with comparable regularization, the excess risk of GD is always within a constant factor of ridge, but ridge can be polynomially worse even when tuned optimally. Second, GD is incomparable with online stochastic gradient descent (SGD). While it is known that for certain problems GD can be polynomially better than SGD, the reverse is also true: we construct problems, inspired by benign overfitting theory, where optimally stopped GD is polynomially worse. Finally, GD dominates SGD for a significant subclass of problems -- those with fast and continuously decaying covariance spectra -- which includes all problems satisfying the standard capacity condition.
This is joint work with Peter Bartlett, Sham Kakade, Jason Lee, and Bin Yu.
- Bio:
Jingfeng Wu is an assistant professor of Statistics at UC Berkeley. His research focuses on deep learning theory, optimization, and statistical learning. He earned his Ph.D. in Computer Science from Johns Hopkins University in 2023. Prior to that, he received a B.S. in Mathematics (2016) and an M.S. in Applied Mathematics (2019), both from Peking University. In 2023, he was recognized as a Rising Star in Data Science by the University of Chicago and UC San Diego.
[Reference]:
- Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, Nathan Srebro. The Implicit Bias of Gradient Descent on Separable Data.
[ arXiv:1710.10345 ]. ICLR 2018. Gradient descent on logistic regression leads to max margin.
- Matus Telgarsky. Margins, Shrinkage, and Boosting. [ arXiv:1303.4172 ]. ICML 2013. An older paper on gradient descent on exponential/logistic loss
leads to max margin.
- Yuan Yao, Lorenzo Rosasco and Andrea Caponnetto. On Early Stopping in Gradient Descent Learning. Constructive Approximation, 2007, 26 (2): 289-315.
[ link ]
- Jingfeng Wu, Peter L. Bartlett, Jason D. Lee, Sham M. Kakade, and Bin Yu. Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization. [ arXiv:2509.17251 ]
|
Y.Y. |
|
| 10/05/2026, Mon |
Lecture 05: Decision Tree, Bagging, Random Forests and Boosting [ YY's slides ]
|
Y.Y. |
|
| 10/12/2026, Mon |
Lecture 06: Support Vector Machines [ YY's slides ]
|
Y.Y. |
|
by YAO, Yuan.