Summary
This is a data systems course: how large-scale data is stored, organized, processed, streamed,
and fed to machine learning and LLMs. We study the designs behind the systems that run the modern data
platform — MapReduce and its successors, distributed file systems and object storage, columnar formats
and the lakehouse, key-value stores and LSM trees, Spark and vectorized analytical engines, Kafka and
stream processing, sketches and sampling, graph processing, the data pipelines behind large language
models, and large-scale experimentation. The course is built around research papers (two per lecture),
hands-on labs on your own laptop, a paper presentation, and a written final exam.
The recurring theme is not reading everything: the best data systems are the ones that touch the
least data to answer a question.
Course Information
- Fall semester 2026
- Class: Friday 12:00–15:00
- Instructor: Alexandros Ntoulas — office hours Friday 11:00–12:00 or by appointment
antoulas@di.uoa.gr (put M111 in the subject line) - Announcements, course material, and submissions on eClass
Announcements
- Week 1: Welcome to M111! Please enroll on eClass — slides, handouts, lab material, and announcements are posted there.
References
The course material comes primarily from research papers (OSDI, SOSP, SIGMOD, VLDB, NSDI, KDD, etc.),
listed per lecture in the schedule below. No textbook is required. For background:
- Designing Data-Intensive Applications, Martin Kleppmann, O’Reilly, 2017.
- Database Internals, Alex Petrov, O’Reilly, 2019.
- CMU 15-445/645 Database Systems (lecture videos) for a refresher on DBMS fundamentals.
Syllabus & Schedule
| Lecture | Topic | Papers |
|---|---|---|
| MODULE 1: Foundations | ||
| Lecture 1 | What “Big Data” means now: from MapReduce to the modern data platform | |
| MODULE 2: Storage | ||
| Lecture 2 | Storage: GFS, object stores, Parquet, and the lakehouse |
|
| Lecture 3 | Key-value and wide-column stores: Bigtable, LSM trees, Dynamo |
|
| MODULE 3: Processing | ||
| Lecture 4 | Distributed processing: Spark and the dataflow model | |
| Lecture 5 | Analytical engines: cloud warehouses and the single-node renaissance |
|
| Lecture 6 | Streaming: logs, Kafka, and the dataflow model | |
| MODULE 4: Not Reading Everything | ||
| Lecture 7 | Sketches and sampling: Bloom filters, HyperLogLog, MinHash/LSH | |
| Lecture 8 | Graphs at scale + LLMs as data engineers |
|
| MODULE 5: Big Data Meets AI | ||
| Lecture 9 | Data infrastructure for ML and LLMs | |
| Lecture 10 | Experimentation at scale, privacy, and course synthesis | |
| Lecture 11 | Paper presentations | Schedule posted after sign-up |
| Lecture 12 | Paper presentations | Schedule posted after sign-up |
| Lecture 13 | Final exam | Date TBA |