| Term | Fall 2026 | Instructor | Christopher De Sa |
| Course website | www.cs.cornell.edu/courses/cs4787/2026fa/ | cmd353@cornell.edu | |
| Schedule | MW 7:30–8:45PM | Office hours | Wednesdays 2–3PM |
| Room | Kimball Hall B11 | Office | Gates 426 |
[Canvas] [Discussion]
Description: CS4787 explores the principles behind scalable machine learning systems. The course will cover the algorithmic and the implementation principles that power the current generation of machine learning on big data. We will cover training and inference for both traditional ML algorithms such as linear and logistic regression, as well as deep models such as transformers. Topics will include: estimating statistics of data quickly with subsampling, stochastic gradient descent and other scalable optimization methods, mini-batch training, accelerated methods, adaptive learning rates, methods for scalable deep learning, hyperparameter optimization, parallel and distributed training, quantization and model compression, and efficient inference.
Prerequisites: CS3780 or equivalent, CS 2110 or equivalent
Format: Lectures during the scheduled lecture period will cover the course content. Problem sets will be used to encourage familiarity with the content and develop competence with the more mathematical aspects of the course. Programming assignments will help build intuition and familiarity with how machine learning algorithms run. There will be one midterm exam and one final exam, each of which will test both theoretical knowledge and programmming implementation of concepts.
Material: The course is based on books, papers, and other texts in machine learning, scalable optimization, and systems. Texts will be provided ahead of time on the website on a per-lecture basis. You aren't expected to necessarily read the texts, but they will provide useful background for the material we are discussing.
Grading: Students taking CS4787 will be evaluated on the following basis.
| 15% | Problem sets |
| 35% | Programming assignments |
| 20% | Prelim Exam |
| 30% | Final Exam |
CS5777 has an additional paper-reading component, and students taking CS5777 will be evaluated as follows.
| 10% | Problem sets |
| 30% | Programming assignments |
| 10% | Paper reading |
| 20% | Prelim Exam |
| 30% | Final Exam |
New this year, your knowledge of the content from the problem sets and the programming assignments will also be checked via short oral “mastery checks” with the TAs. Logistical details of how these will be conducted and scheduled are still being finalized and will be announced later in the semester.
Inclusiveness: You should expect and demand to be treated by your classmates and the course staff with respect. You belong here, and we are here to help you learn—and enjoy—this course. If any incident occurs that challenges this commitment to a supportive and inclusive environment, please let the instructor know so that we can address the issue. We are personally committed to this, and subscribe to the Computer Science Department's Values of Inclusion.
AI Use Policy: It is an academic integrity violation to represent the output of a generative AI tool as your own work. You may not submit any work produced by generative AI as part of the solution of problem sets, programming assignments, or paper reading assignments. It is an academic integrity violation to use any sort of generative AI to assist you during an exam.
Beyond this, you may use AI as you like to assist your own learning in the course, including by asking the AI about course content, asking it to document functions, asking it for the name of a function that does something, asking it to explain an error message, etc.
Course calendar may be subject to change.
| Monday, August 24 Aug 23Aug 24Aug 25Aug 26Aug 27Aug 28Aug 29 |
Monday, August 24
Lecture 1. Introduction and course overview. [Notes PDF]
Problem Set 1 Released. [Notebook] [HTML]
|
| Wednesday, August 26 Aug 23Aug 24Aug 25Aug 26Aug 27Aug 28Aug 29 |
Wednesday, August 26
Lecture 2. Linear algebra done efficiently: Mapping mathematics to numpy. ML via efficient kernels linked together in python. [Notebook] [HTML]
Background reading material:
|
| Monday, August 31 Aug 30Aug 31Sep 1Sep 2Sep 3Sep 4Sep 5 |
Monday, August 31
Lecture 3. Software for learning with gradients. Numerical differentiation, symbolic differentiation, and automatic differentiation. Efficient gradients with backpropagation. [Notebook] [HTML]
Background reading material:
|
| Wednesday, September 2 Aug 30Aug 31Sep 1Sep 2Sep 3Sep 4Sep 5 |
Wednesday, September 2
Background reading material:
Problem Set 1 Due.
|
| Monday, September 7 Sep 6Sep 7Sep 8Sep 9Sep 10Sep 11Sep 12 |
Monday, September 7
Labor Day. No Lecture.
|
| Wednesday, September 9 Sep 6Sep 7Sep 8Sep 9Sep 10Sep 11Sep 12 |
Wednesday, September 9
Lecture 5. Scaling to complex models by learning with optimization algorithms. Learning in the underparameterized regime. Gradient descent, convex optimization and conditioning. Stochastic gradient descent. [Notebook] [HTML]
Background reading material:
|
| Monday, September 14 Sep 13Sep 14Sep 15Sep 16Sep 17Sep 18Sep 19 |
Monday, September 14
Lecture 6. Adapting algorithms to hardware. Minibatching and the effect of the learning rate. Our first hyperparameters. [Notebook] [HTML]
Background reading material:
|
| Wednesday, September 16 Sep 13Sep 14Sep 15Sep 16Sep 17Sep 18Sep 19 |
Wednesday, September 16
Lecture 7. Optimization techniques for efficient ML. Accelerating SGD with momentum. [Notebook] [HTML]
Background reading material:
|
| Monday, September 21 Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25Sep 26 |
Monday, September 21
Lecture 8. Optimization techniques for efficient ML, continued. Accelerating SGD with preconditioning and adaptive learning rates. [Notebook] [HTML]
Background reading material:
|
| Wednesday, September 23 Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25Sep 26 |
Wednesday, September 23
Background reading material:
|
| Monday, September 28 Sep 27Sep 28Sep 29Sep 30Oct 1Oct 2Oct 3 |
Monday, September 28
Lecture 10. Deep neural networks review. The overparameterized regime and how it affects optimization. Matrix multiply as computational core of learning. [Notebook] [HTML]
Background reading material:
|
| Wednesday, September 30 Sep 27Sep 28Sep 29Sep 30Oct 1Oct 2Oct 3 |
Wednesday, September 30
Lecture 11. Deep neural networks review continued. Transformers and sequence models. [Notebook] [HTML]
Background reading material:
|
| Monday, October 5 Oct 4Oct 5Oct 6Oct 7Oct 8Oct 9Oct 10 |
Monday, October 5
Lecture 12. Hyperparameter Optimization. Grid search. Random search. Manual hyperparameter tuning. [Notebook] [HTML]
Background reading material:
|
| Wednesday, October 7 Oct 4Oct 5Oct 6Oct 7Oct 8Oct 9Oct 10 |
Wednesday, October 7
|
| Thursday, October 8 Oct 4Oct 5Oct 6Oct 7Oct 8Oct 9Oct 10 |
Thursday, October 8
Prelim Exam. 7:30PM, WRNB25, WRNB75.
|
| Monday, October 12 Oct 11Oct 12Oct 13Oct 14Oct 15Oct 16Oct 17 |
Monday, October 12
Fall Break. No Lecture.
|
| Wednesday, October 14 Oct 11Oct 12Oct 13Oct 14Oct 15Oct 16Oct 17 |
Wednesday, October 14
Lecture 14. Scaling laws.
Background reading material:
|
| Monday, October 19 Oct 18Oct 19Oct 20Oct 21Oct 22Oct 23Oct 24 |
Monday, October 19
Background reading material:
|
| Wednesday, October 21 Oct 18Oct 19Oct 20Oct 21Oct 22Oct 23Oct 24 |
Wednesday, October 21
|
| Monday, October 26 Oct 25Oct 26Oct 27Oct 28Oct 29Oct 30Oct 31 |
Monday, October 26
Lecture 17. Floating-point arithmetic. Quantized, low-precision machine learning. [Notes PDF]
Background reading material:
|
| Wednesday, October 28 Oct 25Oct 26Oct 27Oct 28Oct 29Oct 30Oct 31 |
Wednesday, October 28
Lecture 18. Parallelism on the GPU: Kernels and Warps. [Notes PDF]
|
| Monday, November 2 Nov 1Nov 2Nov 3Nov 4Nov 5Nov 6Nov 7 |
Monday, November 2
Lecture 19. Parallelism on the GPU 2: CUDA and TensorCores and NVLink. [Notebook] [HTML] [Demo Notebook] [Demo HTML]
|
| Wednesday, November 4 Nov 1Nov 2Nov 3Nov 4Nov 5Nov 6Nov 7 |
Wednesday, November 4
Background reading material:
|
| Monday, November 9 Nov 8Nov 9Nov 10Nov 11Nov 12Nov 13Nov 14 |
Monday, November 9
Lecture 21. Distributed learning 2: More sophisticated patterns; fully-sharded data parallel; distributed inference. [Slides PDF]
Background reading material:
|
| Wednesday, November 11 Nov 8Nov 9Nov 10Nov 11Nov 12Nov 13Nov 14 |
Wednesday, November 11
Lecture 22. Machine learning on hardware beyond GPUs. ML Accelerators. [Slides PDF]
Background reading material:
|
| Monday, November 16 Nov 15Nov 16Nov 17Nov 18Nov 19Nov 20Nov 21 |
Monday, November 16
Lecture 23. Deployment and low-latency inference. Real-time learning. Deep neural network compression and pruning. [Slides PDF]
Background reading material:
|
| Wednesday, November 18 Nov 15Nov 16Nov 17Nov 18Nov 19Nov 20Nov 21 |
Wednesday, November 18
Lecture 24. Foundation Models. Transfer Learning. In-context learning. Fine-tuning. [Notes PDF]
|
| Monday, November 23 Nov 22Nov 23Nov 24Nov 25Nov 26Nov 27Nov 28 |
Monday, November 23
Lecture 25. Online learning. Multimodal learning and tokenization. [Notes PDF]
Background reading material:
|
| Wednesday, November 25 Nov 22Nov 23Nov 24Nov 25Nov 26Nov 27Nov 28 |
Wednesday, November 25
Thanksgiving Break. No Lecture.
|
| Monday, November 30 Nov 29Nov 30Dec 1Dec 2Dec 3Dec 4Dec 5 |
Monday, November 30
Lecture 26. Alternatives to Autoregressive Transformers. Diffusion models. State-space models. [Notebook] [HTML]
|
| Wednesday, December 2 Nov 29Nov 30Dec 1Dec 2Dec 3Dec 4Dec 5 |
Wednesday, December 2
|
| Monday, December 7 Dec 6Dec 7Dec 8Dec 9Dec 10Dec 11Dec 12 |
Monday, December 7
|