Software Engineering for Data Scientists
Apply software engineering practices to Python data science projects, to produce robust code that your colleagues can trust.
About the course
This course is for Python data science programmers who need a more efficient and trustworthy development process. Practices from software engineering, covering testing, code design, documentation, debugging and profiling, are distilled into hands-on exercises.
The course assumes a basic level of fluency with Python (built-in data types, control flow statements) and the PyData ecosystem (the basics of pandas).
What you will learn
- Testing and debugging your code, so you can trust it
- Software design and documentation for more maintainable code
- Writing more robust production code, to reduce downtime and other code issues
- Collaborating with other technical team members with more confidence
Syllabus
Fundamentals of software testing
- Python tools: unittest, mock, pytest, coverage, Hypothesis
- Defensive programming vs unit testing vs test-driven development
Structuring Python code
- Notebooks vs scripts vs packages vs modules
- Designing maintainable and reusable code
Effective documentation
- Docstrings and documentation styles
- Python tools: Sphinx
Logging and debugging
- Configuring Python's logging module
- Debugging Python code with pdb
Profiling and optimisation
- Finding bottlenecks in your Python code
Refactoring exercises
- Putting everything together on realistic code