TDT4225
Very Large, Distributed Data Volumes
Autumn
Trondheim
English
About this course
Content
Introduction to large and distributed data volumes. Introduction to distributed techniques. How to design data-intensive applications? Reliability, scalability, and maintainability; how we need to think about them; and how we can achieve them. Data models and query languages. Indexing and storage techniques. Encoding of data. Replication, partitioning and transactions. Fault models, consistency and consensus.
Learning outcomes
Learning outcome
Knowledge:
By completion of this course, the candidate should be able to explain
1. reliable, scalable, and maintainable distributed systems
2. data models and query languages
3. indexing- and data storage methods
4. formats for encoding of data
5. models of replication
6. models of partitioning
7. theory of transactions and concurrency
8. fault models
9. consistency and consensus
10. synchronization of clocks
11. distributed debugging
12. content delivery networks and dirstibuted hash tables
13. database-as-a-service
14. trustworthy data systems
Skills:
By completion of this course, the candidate should be able to
1. develop applications with big data using standard database products.
2. evaluate existing systems and solutions for distributed storage and management of data
3. combine tools to build the properties you need
4. develop new systems for distributed storing and management of data
General competence:
By completion of this course, the student should be able to explain distributed systems.
Teaching methods
Lectures, exercises, projects and self-study.
There are compulsory exercises in the subject.
There are two projects involving programming with big data volumes and which are done in small groups. Each of these counts for 20% of the final grade. Total: 40%.
The final exam counts for 60% of the final grade.