Plain-English companion pages for Daniel Jurafsky & James H. Martin — Speech and Language Processing, 3rd edition. These notes focus on the mid-semester NLP arc: statistical language modeling, vector meaning, neural networks for language, and sequence labeling with POS tagging, HMMs, Viterbi, and modern neural taggers.
Probability of sequences, chain rule, Markov assumptions, MLE, log probabilities, generation, evaluation, perplexity, smoothing, backoff, interpolation, Kneser-Ney, and unknown words.
chapter 6Lexical semantics, distributional hypothesis, term-document and word-context matrices, TF-IDF, PPMI, cosine, Word2Vec, CBOW, GloVe, evaluation, and bias.
chapter 7Weighted units, bias, activations, XOR, depth, feedforward networks, softmax, loss, gradient descent, backpropagation, and neural language models.
chapter 8POS tagsets, ambiguity, HMM transition A, emission B, Viterbi, forward algorithm, MEMMs, bidirectionality, and Bi-LSTM-CRF.