CS388: Natural Language Processing (Fall 2026)
Instructor: Elias Stengel-Eskin, esteng@utexas.edu
Lecture: Tuesday and Thursday 3:30pm - 5:00pm, JGB 2.202
Instructor Office Hours: Tuesdays, 2:30-2:30pm in GDC 3.810
TA: Ashwin Vinod
TA Office Hours: Monday, 2-3pm (zoom); Thursday 2-3pm (in-person).
Unique Number: 55645
Description
This class is a graduate-level introduction to Natural Language Processing (NLP), the study of computing systems that
can process, understand, or communicate in human language. The last several years have reshaped the field: a single class of
models — large language models (LLMs) pre-trained on text and then adapted through supervised learning and reinforcement
learning — now underpins most NLP applications. This course is about understanding how those systems work, why they work,
where they fail, and what we still do not know about them.
The course builds the modeling toolkit from the ground up: text classification and feature extraction, n-gram language models,
word embeddings, feedforward neural networks, and then the Transformer architecture and the encoder and decoder models built on it.
We treat language modeling as the unifying thread — generation as V-way classification — and trace it through pre-training,
mid-training, and post-training, including reinforcement learning from human feedback (RLHF) and from verifiable rewards (RLVR).
The second half covers what these models are used for and what is known about their behavior: in-context learning, chain-of-thought
reasoning, dataset artifacts and bias, multimodal grounding and vision-language models, agents, and applications spanning
translation, code, efficiency, interpretability, and safety.
Throughout, the emphasis is on connecting course material to the current state of the art. You should leave this course able to
read a current NLP paper critically, understand the design decisions behind a modern LLM, and carry out original research of your own.
Requirements
- 391L - Machine Learning, 343 - Artificial Intelligence, or equivalent AI/ML course experience
- Familiarity with Python for homeworks and the final project
- Additional prior exposure to discrete math, probability, linear algebra, optimization, linguistics, and NLP
useful but not required
Assessment: 4 quizzes (20%, lowest dropped), 4 homeworks (10%, lowest dropped), midterm (15%), final exam (30%),
final project (25%). Quizzes and exams are in person and closed-book; quizzes are given in the first 15 minutes of class on the
dates below. Note that the quiz dates below are subject to change depending on course progress. The final exam is given
in class on the final day of class (Dec 3). See the syllabus for full details, including the AI use policy, which is
permissive on homeworks and carries specific obligations on the final project.
Assignments: See syllabus for more details about these.
Homework 1: Classification and Neural Networks [to be posted]
Homework 2: Language Modeling and Transformers [to be posted]
Homework 3: Post-training and RL [to be posted]
Homework 4: Applications [to be posted]
Final Project [to be posted]
Readings: Textbook readings are assigned to complement the material discussed in lecture. You may find it useful
to do these readings before lecture as preparation or after lecture to review, but you are not expected to know everything discussed
in the textbook if it isn't covered in lecture.
Paper readings are intended to supplement the course material if you are interested in diving deeper on particular topics.
Readings are not required for the quizzes or exams unless the material was also covered in lecture.
The chief text in this course is Jurafsky and Martin: Speech and Language Processing (3rd ed.),
available as a free PDF online; its recent chapters cover Transformers and LLMs. This is supplemented by
Eisenstein: Natural Language Processing
for classification and structured prediction, and
Goldberg: A Primer on Neural Network Models for Natural Language Processing for
neural network fundamentals. Much of the second half of the course has no textbook treatment and is covered by papers.
Slides will be linked from this table as the semester progresses.
| Date |
Topics |
Readings |
Assignments |
| Aug 25 |
Introduction: NLP Tasks, Ambiguity, and a Brief History of Modern NLP |
JM 2
|
|
| Aug 27 |
Machine Learning Basics; Binary Classification |
Eisenstein 2.0-2.5, 4.2-4.4.1
JM 4
JM 5
Duchi+11 AdaGrad
|
HW1 out |
| Sep 1 |
Multi-class Classification; Feature Extraction; N-gram Language Models, Smoothing and Backoff; Neural Net History |
Eisenstein 4.2
JM 5.3-5.6
JM 3 (N-gram LMs)
Neural net history (optional — background only, not required):
HochreiterSchmidhuber97 LSTMs
LeCun+98 Convnets
Henderson03 Neural Parsing
Collobert+11 NLP (Almost) from Scratch
Krizhevsky+12 AlexNet
Socher+13 Recursive Models / Sentiment Treebank
ChenManning14 Dependency Parsing
Kim14 CNNs for Sentence Classification
Kalchbrenner+14 CNNs for Modelling Sentences
Sutskever+14 Seq2seq
Bahdanau+14 Attention for NMT
|
|
| Sep 3 |
Neural Networks: Feedforward, Backpropagation |
Eisenstein 3.0-3.3
JM 7
Goldberg 4
Bengio+03 NPLM
MinskyPapert69 Perceptrons (XOR)
KingmaBa15 Adam
Srivastava+14 Dropout
IoffeSzegedy15 Batch Normalization
Olah, Neural Networks, Manifolds, and Topology
|
HW1 due |
| Sep 8 |
Word Embeddings; Bias in Embeddings |
JM 6
Mikolov+13 word2vec
Pennington+14 GloVe
LevyGoldberg+15 Improving Similarity
Bolukbasi+16 Gender
|
Quiz 1 |
| Sep 10 |
Neural Language Models, RNNs, and Attention; Positional Encodings |
Bengio+03 NPLM
Elman90 Finding Structure in Time
Luong+15 Attention
Bahdanau+14 Attention for NMT
Alammar Illustrated Transformer
Su+21 RoPE
Kazemnejad+23 NoPE / Positional Encoding and Length Generalization
Raschka, No Positional Embeddings (NoPE)
|
|
| Sep 15 |
Transformers 1: Self-Attention, Architecture |
Vaswani+17 Transformers
JM 9
Alammar Illustrated Transformer
PhuongHutter Formal Algorithms
|
|
| Sep 17 |
Transformers 2: Positional Encoding, Scaling |
Su+21 RoPE
Kaplan+20 Scaling Laws
Hoffmann+22 Chinchilla
ZhangSennrich19 RMSNorm
|
HW2 out |
| Sep 22 |
Encoders: BERT, Tokenization |
Devlin+19 BERT
Alammar Illustrated BERT
Liu+19 RoBERTa
Clark+20 ELECTRA
Sennrich+16 BPE
BostromDurrett20 Tokenizers
|
|
| Sep 24 |
Decoders: GPT/T5, Decoding Methods |
Radford+19 GPT-2
Brown+20 GPT-3
Raffel+19 T5
Lewis+19 BART
Holtzman+19 Nucleus Sampling
|
HW2 due |
| Sep 29 |
Evaluation, Datasets, and Dataset Artifacts |
Wang+19 SuperGLUE
Gururangan+18 Artifacts
McCoy+19 HANS
Gardner+20 Contrast Sets
Swayamdipta+20 Cartography
|
Quiz 2 |
| Oct 1 |
In-Context Learning |
Brown+20 GPT-3
Zhao+21 Calibrate Before Use
Min+22 Rethinking Demonstrations
Xie+21 ICL as Implicit Bayesian Inference
Olsson+22 Induction Heads
|
|
| Oct 6 |
Chain-of-Thought Reasoning |
Wei+22 CoT
Kojima+22 Step-by-step
Wang+22 Self-Consistency
Gao+22 PAL
Turpin+23 Unfaithful CoT
|
|
| Oct 8 |
MIDTERM EXAM (in class) |
|
|
| Oct 13 |
Building an LLM 1: Pre-training and Mid-training |
Hoffmann+22 Chinchilla
Touvron+23 Llama 2
Groeneveld+24 OLMo
Soldaini+24 Dolma
|
|
| Oct 15 |
Reinforcement Learning for NLP |
SuttonBarto 3, 13
Schulman+17 PPO
Ramamurthy+22 RL4LMs
|
HW3 out |
| Oct 20 |
RLHF |
Stiennon+20 Learning to Summarize
Ouyang+22 InstructGPT
Bai+22 Constitutional AI
Rafailov+23 DPO
Singhal+23 Length
|
|
| Oct 22 |
RLVR and Reasoning Models |
DeepSeek-AI+25 DeepSeek-R1
Lambert+24 Tulu 3
Shao+24 GRPO / DeepSeekMath
Lightman+23 Process Supervision
|
HW3 due FP proposal due Oct 23 |
| Oct 27 |
Multimodal Grounding 1: Intro to Vision, CLIP |
Dosovitskiy+20 ViT
Radford+21 CLIP
He+15 ResNet
|
Quiz 3 |
| Oct 29 |
Multimodal Grounding 2: VLMs and VLAs |
Alayrac+22 Flamingo
Liu+23 LLaVA
Driess+23 PaLM-E
Brohan+23 RT-2
|
|
| Nov 3 |
Agents |
Yao+22 ReAct
Schick+23 Toolformer
Yao+23 Tree of Thoughts
Jimenez+23 SWE-bench
Yao+22 WebShop
|
HW4 out |
| Nov 5 |
Applications: Machine Translation and Multilinguality |
Eisenstein 18.1-18.2, 18.4
Liu+20 mBART
NLLB+22 No Language Left Behind
Conneau+19 XLM-R
Kocmi+23 LLMs for MT Eval
|
|
| Nov 10 |
Applications: Code (Semantic Parsing, Text-to-Code, Code Agents) |
ZettlemoyerCollins05 Semantic Parsing
Berant+13 Freebase QA
Chen+21 Codex
Austin+21 Program Synthesis
Jimenez+23 SWE-bench
|
|
| Nov 12 |
Applications: Efficiency |
Hu+21 LoRA
Dettmers+23 QLoRA
Dao+22 FlashAttention
Leviathan+23 Speculative Decoding
Kwon+23 vLLM / PagedAttention
|
HW4 due |
| Nov 17 |
Applications: Interpretability |
Lipton16 Mythos
Ribeiro+16 LIME
Sundararajan+17 Integrated Gradients
Meng+22 ROME
Bricken+23 Monosemanticity
|
Quiz 4 |
| Nov 19 |
Applications: Safety |
Zou+23 Universal Attacks
Shen+23 Jailbreaking
Ganguli+22 Red Teaming
BenderGebru+21 Stochastic Parrots
|
|
| Nov 24 |
No class — Thanksgiving |
|
|
| Nov 26 |
No class — Thanksgiving |
|
|
| Dec 1 |
Wrap-up and Review |
|
|
| Dec 3 |
FINAL EXAM (in class, final day of class) |
|
FP report due Dec 11 |