Data as the foundation of ML systems -- representation, data models, storage layouts, serialization, data life cycle, architecture patterns, modern data stacks, pipelines, storage infrastructure, serving, DataOps, observability, and continuous training/deployment. This vault mirrors the NLP subject vault: cheatsheet, slides explained, question bank, formula sheet, book guide, and references. Covers CS1 through L5 for DSE/AIML ZG529.
Dense one-glance reference -- data formats, models, storage layouts, ETL/ELT, warehouse/lake/lakehouse, Lambda/Kappa, pipelines, DataOps and ML infrastructure.
02 · slides explainedEvery uploaded deck unpacked in plain easy words: CS1 Data Representations, CS2 fundamentals, L3 architectures, L4 pipelines, and L5 infrastructure/DataOps.
03 · question bankExam-style questions pulled from the decks. Questions only -- no answers -- grouped lecture by lecture for active recall.
04 · formula sheetOperational formulas and decision tables for reliability, freshness, sampling, throughput, storage cost, retry backoff, availability and data quality.
05 · book guideA study map across Fundamentals of Data Engineering, DDIA, Reliable Machine Learning, Data Pipelines Pocket Reference and Effective Data Science Infrastructure.
06 · referencesBooks, course-aligned readings, key architecture terms, and practical tools for data management and ML data infrastructure.
01 · CHEATSHEET
A compact revision card for the course arc: data as raw material and liability, structured/semi-structured/unstructured data, relational/document/graph/key-value models, row vs column layout, serialization formats, OLTP vs OLAP, data platform components, architecture patterns, pipelines and DataOps.
02 · SLIDES EXPLAINED
The root Slides/ decks rewritten as concept notes: why data management matters for ML, how source data becomes usable training and serving data, where warehouses/lakes/lakehouses fit, why Lambda and Kappa exist, how batch/stream/CDC pipelines move data, and how DataOps keeps the system trustworthy.
03 · QUESTION BANK
Questions extracted from CS1 through L5 -- questions only, no answers -- for recall practice before checking the slides explained and cheatsheet. Includes short answers, comparison prompts, scenario questions and design exercises.
04 · FORMULA SHEET
DMML is mostly architecture and systems, but exams still reward precise operational reasoning. This sheet collects the useful formulas: availability, reliability, freshness, throughput, latency, sampling, storage cost, compression ratio, CDC lag, retry backoff and data-quality rates.
05 · BOOK GUIDE
A book companion for the references named inside the decks: Fundamentals of Data Engineering, Designing Data-Intensive Applications, Reliable Machine Learning, Data Pipelines Pocket Reference, and Effective Data Science Infrastructure -- mapped to the lectures.
06 · REFERENCES
The core books, architecture patterns, practical tools and external readings worth keeping: warehouses/lakes/lakehouses, ETL/ELT, CDC, streaming, orchestration, IaC/GitOps, data quality, observability and MLOps.