Computational techniques for large-scale data
About
The aim of this course is to deepen the students’ knowledge and skills and familiarize them with the technical and technological side of data science, including software respectively hardware environments. The course will introduce aspects of designing and implementing large-scale data science solutions.
In particular, the course will include:
- an overview of computer architectures, algorithmic approaches, and high- performance computing infrastructures with a focus on limitations for processing large-scale data,
- an introduction to relevant frameworks for cluster computing with large-scale data,
- implementation of data analysis tools on a cluster using Python and appropriate software frameworks,
- data structures and algorithms, such as index structures, which can greatly accelerate computations with large-scale data
Prerequisites and selection
Entry requirements
To be eligible to the course, the student should have a Bachelor's degree in any subject, or have successfully completed 90 credits of studies in computer science, software engineering, or equivalent. Specifically, at least 15 credits of successfully completed courses in programming, or equivalent are required. The student needs to have successfully completed a course in probability theory or statistics.
Applicants must prove knowledge of English: English 6/English level 2 or the equivalent level of an internationally recognized test, for example TOEFL, IELTS.
Selection
Selection is based upon the number of credits from previous university studies, maximum 285 credits