← Books Explained

// BOOK COMPANION · 25 OF 25 LIVE

Speech and Language Processing — explained.

An original, chapter-by-chapter companion to Speech and Language Processing by Daniel Jurafsky and James H. Martin, 3rd edition draft. Written as a first-principles story for master's students: every finished chapter is at least 3,000 words with the conceptual map, derivation, worked case, evaluation design, failure modes, study lab, oral-exam questions, and a complete summary.

Aligned to the official January 6, 2026 online draft. These notes do not copy or replace the book; they are independent educational explanations. Read the official draft ↗


VOLUME I · LARGE LANGUAGE MODELS

chapter 1

Introduction live

The map of language technology: what NLP and speech systems do, why language is difficult, and how the book connects models, data, linguistic structure, and evaluation.

4,836 words· 25 min
chapter 2

Words and Tokens live

Words, morphemes, Unicode, subword tokenization, corpora, regular expressions, rule-based tokenization, and minimum edit distance.

4,702 words· 24 min
chapter 3

N-gram Language Models live

Conditional probability, n-grams, training and test splits, perplexity, sampling, overfitting, smoothing, interpolation, backoff, and entropy.

4,640 words· 24 min
chapter 4

Logistic Regression and Text Classification live

Supervised classification, the sigmoid and softmax, features, cross-entropy, gradient descent, precision, recall, F1, cross-validation, and significance testing.

4,611 words· 24 min
chapter 5

Embeddings live

Lexical semantics, distributional meaning, count vectors, cosine similarity, word2vec, semantic properties, bias, and evaluation.

4,544 words· 23 min
chapter 6

Neural Networks live

Units, nonlinear activation, XOR, feedforward networks, classification, embedding inputs, backpropagation, optimization, and regularization.

4,490 words· 23 min
chapter 7

Large Language Models live

Language-model architectures, conditional generation, prompting, decoding, pretraining, scaling, evaluation, ethical risk, and safety.

4,580 words· 23 min
chapter 8

Transformers live

Self-attention, transformer blocks, parallel computation, token and positional embeddings, language-model heads, sampling, training, scaling, and interpretation.

4,557 words· 23 min
chapter 9

Post-training: Instruction Tuning, Alignment, and Test-Time Compute live

Supervised instruction tuning, preference data, reward modeling, RLHF-style optimization, direct preference optimization, alignment limits, reasoning, and test-time compute.

4,652 words· 24 min
chapter 10

Masked Language Models live

Bidirectional transformer encoders, masked-token training, contextual embeddings, fine-tuning for classification, and sequence labeling.

4,545 words· 23 min
chapter 11

Information Retrieval and Retrieval-Augmented Generation live

Sparse and dense retrieval, inverted indexes, ranking, evaluation, question answering, retrieval-augmented generation, grounding, and datasets.

4,630 words· 24 min
chapter 12

Machine Translation live

Cross-language divergence, encoder-decoder translation, attention, beam search, low-resource translation, evaluation, and bias.

4,542 words· 23 min
chapter 13

RNNs and LSTMs live

Recurrent neural networks, sequence modeling, stacked and bidirectional architectures, long short-term memory, encoder-decoder models, and attention.

4,561 words· 23 min
chapter 14

Phonetics and Speech Feature Extraction live

Speech sounds, phonetic transcription, articulation, prosody, acoustic signals, spectrograms, log-Mel features, and MFCCs.

4,557 words· 23 min
chapter 15

Automatic Speech Recognition live

The ASR task, convolutional encoders, encoder-decoder recognition, self-supervised speech models, CTC, decoding, and word error rate.

4,593 words· 23 min
chapter 16

Text-to-Speech live

TTS pipelines, audio codecs and discrete tokens, language-model speech generation, VALL-E-style two-stage systems, evaluation, other speech tasks, and spoken language models.

4,543 words· 23 min

VOLUME II · ANNOTATING LINGUISTIC STRUCTURE

chapter 17

Sequence Labeling for Parts of Speech and Named Entities live

Word classes, part-of-speech tagging, named entities, HMMs, conditional random fields, neural sequence labeling, and span-level evaluation.

4,560 words· 23 min
chapter 18

Context-Free Grammars and Constituency Parsing live

Constituents, context-free grammars, treebanks, ambiguity, normal forms, CKY dynamic programming, neural span parsing, and evaluation.

4,561 words· 23 min
chapter 19

Dependency Parsing live

Head-dependent relations, dependency trees, transition-based parsing, graph-based parsing, projectivity, decoding, and attachment evaluation.

4,493 words· 23 min
chapter 20

Information Extraction: Relations, Events, and Time live

Relation extraction, event extraction, temporal representation, aspect, TimeBank-style annotation, temporal analysis, and template filling.

4,558 words· 23 min
chapter 21

Semantic Role Labeling and Argument Structure live

Semantic roles, alternations, thematic-role limits, PropBank, FrameNet, SRL systems, selectional restrictions, and predicate decomposition.

4,551 words· 23 min
chapter 22

Lexicons for Sentiment, Affect, and Connotation live

Emotion definitions, sentiment and affect lexicons, human labeling, semi-supervised induction, supervised word sentiment, lexicon-based recognition, entity-centered affect, and connotation frames.

4,574 words· 23 min
chapter 23

Coreference Resolution and Entity Linking live

Mentions, anaphora, coreference clusters, mention detection, ranking architectures, entity linking, evaluation, Winograd-style cases, and gender bias.

4,546 words· 23 min
chapter 24

Discourse Coherence live

Coherence relations, discourse structure, centering, entity-based coherence, local representation learning, global coherence, and argument structure.

4,496 words· 23 min
chapter 25

Conversation and Its Structure live

Turn-taking, adjacency pairs, grounding, repair, dialog acts, conversational corpora, context, initiative, and evaluation.

4,544 words· 23 min