Summary

This is a data systems course: how large-scale data is stored, organized, processed, streamed, and fed to machine learning and LLMs. We study the designs behind the systems that run the modern data platform — MapReduce and its successors, distributed file systems and object storage, columnar formats and the lakehouse, key-value stores and LSM trees, Spark and vectorized analytical engines, Kafka and stream processing, sketches and sampling, graph processing, the data pipelines behind large language models, and large-scale experimentation. The course is built around research papers (two per lecture), hands-on labs on your own laptop, a paper presentation, and a written final exam. The recurring theme is not reading everything: the best data systems are the ones that touch the least data to answer a question.

Course Information

Announcements

References

The course material comes primarily from research papers (OSDI, SOSP, SIGMOD, VLDB, NSDI, KDD, etc.), listed per lecture in the schedule below. No textbook is required. For background:

Syllabus & Schedule

LectureTopicPapers
MODULE 1: Foundations
Lecture 1What “Big Data” means now: from MapReduce to the modern data platform
  • MapReduce: Simplified Data Processing on Large Clusters, J. Dean, S. Ghemawat, OSDI 2004 [pdf]
  • What Goes Around Comes Around... And Around, M. Stonebraker, A. Pavlo, SIGMOD Record 2024 [pdf]
  • Lab 0 out (setup: Python + DuckDB; details on eClass)
MODULE 2: Storage
Lecture 2Storage: GFS, object stores, Parquet, and the lakehouse
  • The Google File System, S. Ghemawat, H. Gobioff, S.-T. Leung, SOSP 2003 [pdf]
  • Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics, M. Armbrust et al., CIDR 2021 [pdf]
  • Optional: BtrBlocks: Efficient Columnar Compression for Data Lakes, M. Kuschewski et al., SIGMOD 2023 [pdf]
  • Lab 1 out (on eClass)
Lecture 3Key-value and wide-column stores: Bigtable, LSM trees, Dynamo
  • Bigtable: A Distributed Storage System for Structured Data, F. Chang et al., OSDI 2006 [pdf]
  • Dynamo: Amazon’s Highly Available Key-value Store, G. DeCandia et al., SOSP 2007 [pdf]
  • Optional: Evolution of Development Priorities in Key-value Stores Serving Large-scale Applications: The RocksDB Experience, S. Dong et al., FAST 2020 [page]
  • Presentation paper list posted
MODULE 3: Processing
Lecture 4Distributed processing: Spark and the dataflow model
  • Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing, M. Zaharia et al., NSDI 2012 [pdf]
  • Scalability! But at what COST?, F. McSherry, M. Isard, D. Murray, HotOS 2015 [pdf]
  • Presentation sign-up closes
Lecture 5Analytical engines: cloud warehouses and the single-node renaissance
  • The Snowflake Elastic Data Warehouse, B. Dageville et al., SIGMOD 2016 [pdf]
  • DuckDB: an Embeddable Analytical Database, M. Raasveldt, H. Mühleisen, SIGMOD 2019 [pdf]
  • Optional: Dremel: A Decade of Interactive SQL Analysis at Web Scale, S. Melnik et al., VLDB 2020 [pdf]
  • Lab 1 due (Friday 23:59, eClass)
Lecture 6Streaming: logs, Kafka, and the dataflow model
  • Kafka: a Distributed Messaging System for Log Processing, J. Kreps, N. Narkhede, J. Rao, NetDB 2011 [pdf]
  • The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing, T. Akidau et al., VLDB 2015 [pdf]
MODULE 4: Not Reading Everything
Lecture 7Sketches and sampling: Bloom filters, HyperLogLog, MinHash/LSH
  • HyperLogLog in Practice: Algorithmic Engineering of a State of The Art Cardinality Estimation Algorithm, S. Heule, M. Nunkesser, A. Hall, EDBT 2013 [pdf]
  • Similarity Search in High Dimensions via Hashing, A. Gionis, P. Indyk, R. Motwani, VLDB 1999 [pdf]
  • Lab 2 out (on eClass)
Lecture 8Graphs at scale + LLMs as data engineers
  • Pregel: A System for Large-Scale Graph Processing, G. Malewicz et al., SIGMOD 2010 [pdf]
  • Semantic Operators: A Declarative Model for Rich, AI-based Data Processing (LOTUS), L. Patel et al., 2024–25 [pdf]
  • Optional: One Trillion Edges: Graph Processing at Facebook-Scale, A. Ching et al., VLDB 2015 [pdf] · Can LLM Already Serve as a Database Interface? (BIRD), J. Li et al., NeurIPS 2023 [pdf]
MODULE 5: Big Data Meets AI
Lecture 9Data infrastructure for ML and LLMs
  • The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, G. Penedo et al., NeurIPS 2024 D&B [pdf]
  • Deduplicating Training Data Makes Language Models Better, K. Lee et al., ACL 2022 [pdf]
Lecture 10Experimentation at scale, privacy, and course synthesis
  • Online Controlled Experiments at Large Scale, R. Kohavi et al., KDD 2013 [pdf]
  • Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained, R. Kohavi et al., KDD 2012 [pdf]
  • Lab 2 due (Friday 23:59, eClass) · sample exam questions published
Lecture 11Paper presentationsSchedule posted after sign-up
Lecture 12Paper presentationsSchedule posted after sign-up
Lecture 13Final examDate TBA