Python for Data Analytics
Get familiar with the Python stack for data science, and automate the acquisition, cleaning, processing and analysis of data.
About the course
This course helps analysts, researchers, BI experts, data scientists and developers become fluent in Python, so they can automate the acquisition, cleaning and analysis of data.
The first day is an introduction to the core concepts of Python programming, giving participants the foundations to work independently and flexibly. The next two days focus on using Python tools for data wrangling and analysis. Throughout, the course uses interactive examples and hands-on exercises.
What you will learn
- Running Python code with Jupyter and Anaconda
- Core concepts of the Python language
- Loading, filtering, sorting, transforming, aggregating, analysing and plotting data with pandas
- Mathematical functions and array operations with NumPy
- Visualising data with matplotlib and plotly
Syllabus
Environment set-up
- The Anaconda distribution as a Python data science platform
- Overview of Python virtual environment set-up
- Running code in Jupyter notebooks
Python core concepts
- Built-in data types in Python
- Control flow statements
- Defining and using custom functions
- Working with dates and times
- Accessing data on file (CSV, JSON, …)
Python data science libraries
- pandas
- Working with table-like data in pandas
- Loading data from file into DataFrame objects
- Data transformation and indexing
- Summary statistics over DataFrame objects
- Data aggregation queries (the
groupby()method) - Exploratory analysis of new datasets
- Data visualisation over DataFrames
- Join/merge operations with DataFrames
- Time series operations in pandas
- Working with text data in DataFrames
- NumPy
- Working with NumPy arrays
- Essential operations with NumPy arrays
- Statistics and linear algebra with NumPy
- matplotlib vs plotly for data visualisation