Training

Software Engineering for Data Scientists

Apply software engineering practices to Python data science projects, to produce robust code that your colleagues can trust.

About the course

This course is for Python data science programmers who need a more efficient and trustworthy development process. Practices from software engineering, covering testing, code design, documentation, debugging and profiling, are distilled into hands-on exercises.

The course assumes a basic level of fluency with Python (built-in data types, control flow statements) and the PyData ecosystem (the basics of pandas).

What you will learn

  • Testing and debugging your code, so you can trust it
  • Software design and documentation for more maintainable code
  • Writing more robust production code, to reduce downtime and other code issues
  • Collaborating with other technical team members with more confidence

Syllabus

Fundamentals of software testing

  • Python tools: unittest, mock, pytest, coverage, Hypothesis
  • Defensive programming vs unit testing vs test-driven development

Structuring Python code

  • Notebooks vs scripts vs packages vs modules
  • Designing maintainable and reusable code

Effective documentation

  • Docstrings and documentation styles
  • Python tools: Sphinx

Logging and debugging

  • Configuring Python's logging module
  • Debugging Python code with pdb

Profiling and optimisation

  • Finding bottlenecks in your Python code

Refactoring exercises

  • Putting everything together on realistic code

Contact

To run this course for your team, or to discuss a tailored version, email me with a few lines about the group and what you'd like them to get out of it.

info@bonzaniniconsulting.com