Natural Language Processing
An introduction to the techniques and Python tools for practitioners getting started with natural language processing.
About the course
The course balances theoretical foundations with practical examples in Python. No prior experience with libraries such as NLTK or scikit-learn is required. Experience with Python is very useful but not essential: people who use other languages and tools (Java, C++, C#, JavaScript, MATLAB, Excel or R) will also get a lot out of it.
What you will learn
- Data representations for working effectively with text
- Exploratory techniques to quickly gain insights from text data
- Machine learning techniques to organise documents into categories
- Evaluating the quality of your models
- Ideas for advanced applications using natural language data
Syllabus
NLP foundations
The basic tools and techniques to get started with natural language processing.
- NLP applications and the Python NLP ecosystem: NLTK, spaCy, Gensim, scikit-learn
- Working with text: tokenisation, text pre-processing, regular expressions
- Word frequencies and co-occurrences: stop-words and Zipf's law, mining topics of interest with co-occurrences
- Text representation: n-grams, bag-of-words, word embeddings and document embeddings
Topic modelling
Understanding a document, or a collection of documents, with techniques that go beyond word frequencies.
- A bird's-eye view of a document or a dataset
- Navigating topics and sub-topics in a document or a dataset
Text classification
Classifying documents into a set of predefined categories.
- Categorising documents
- Topic classification
- Sentiment analysis
- Model evaluation: assessing classification quality
- Model introspection: explaining classification results
Advanced applications
An outlook on more advanced NLP problems, so participants leave with ideas and techniques for specific applications, such as:
- Named entity recognition: identifying named entities in text
- Text summarisation: extracting the most useful sentences from one or more documents
- Search engines: retrieving relevant documents from a custom text collection