AI in Motion

Natural Language ProcessingDeep diveIntermediate10:51 video32 chapters

Text Classification from Start to Finish — lecture notes

Build a real text classifier: framing and labelling, tokenisation and TF-IDF features, Naive Bayes and logistic-regression baselines, cross-validation and metrics, imbalance, fine-tuning transformers, zero-shot LLMs, error analysis and monitoring.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Text Classification from Start to Finish

Spam filters, sentiment analysis, routing support tickets, flagging harmful content, detecting fake reviews: text classification is the workhorse of practical NLP. In this deep dive we build a classifier end to end, from collecting labels to choosing between classic models, fine tuned transformers and large language models.

0:202. The task

The task — Text Classification from Start to Finish

Text classification assigns one or more labels from a fixed set to a piece of text. A support ticket might be routed to billing, delivery, login or other. A review might be positive or negative. When a text can have several labels at once, it is called multi label classification.

0:413. The pipeline

The pipeline — Text Classification from Start to Finish

Every project follows the same pipeline. Frame the problem and define the labels. Collect and label examples. Represent text as numbers. Build a simple baseline model, then try stronger ones. Evaluate carefully, reading the errors, and finally deploy and monitor, because language keeps changing.

1:004. Labelling

Labelling — Text Classification from Start to Finish

Most of the quality comes from the labels. Write clear definitions with examples for every label, including tricky edge cases, and have several people label a shared sample to measure agreement. Is refund not received a billing issue or a delivery issue? Decide, and write it down.

1:205. Pause and think

Pause and think — Text Classification from Start to Finish

Pause and think. Two annotators label well, that was something, as positive and negative. What does their disagreement tell you? The text is ambiguous, or the guidelines are unclear. Add guidance, or a mixed label. A model can never be more consistent than the labels it learns from.

1:406. Tokenisation

Tokenisation — Text Classification from Start to Finish

Before modelling, text is split into tokens. Word level tokens give short sequences but a huge vocabulary with unknown words. Character level needs a tiny vocabulary but very long sequences. Sub word tokenisation, used by modern models, keeps common words whole and splits rare ones into meaningful pieces.

2:007. Preprocessing

Preprocessing — Text Classification from Start to Finish

Classic models benefit from light preprocessing: lowercasing, removing noise such as HTML tags and URLs, and handling numbers and emojis consistently. Be careful not to remove meaning. Deleting stop words can remove not, flipping the sentiment. Transformers need very little preprocessing beyond their own tokenizer.

2:208. TF-IDF features

TF-IDF features — Text Classification from Start to Finish

A classic representation is TF IDF. Each document becomes a vector of word weights, where words frequent in this document but rare elsewhere score highest, and ubiquitous words like the get zero. The result is a long, sparse vector that works surprisingly well with linear models.

2:399. n-grams

n-grams — Text Classification from Start to Finish

Adding word pairs, called bigrams, as extra features captures short phrases. Not good and good become different features, so a linear model can learn that the phrase is negative even though good alone is positive. The vocabulary grows, but regularisation keeps the model in check.

2:5910. Multi-label problems

Multi-label problems — Text Classification from Start to Finish

Check whether your problem is multi class or multi label. In multi class problems each text has exactly one label, and a softmax picks among them. In multi label problems a text can have several, such as a news article about both politics and the economy, so each label gets its own yes or no decision.

3:2211. Naive Bayes

Naive Bayes — Text Classification from Start to Finish

The simplest strong baseline is Naive Bayes. It assumes that, given the class, words appear independently, which is obviously false, yet it works surprisingly well. It combines each word’s evidence using Bayes’ theorem, trains in seconds, and was the algorithm behind early spam filters.

3:4112. The Naive Bayes rule

The Naive Bayes rule — Text Classification from Start to Finish

The rule multiplies the prior probability of each class by the probability of every word given that class, and picks the class with the highest score. In practice we add log probabilities to avoid underflow, and smooth the counts so an unseen word does not zero out a class.

4:0213. Pause and think

Pause and think — Text Classification from Start to Finish

Pause and think. The word refund never appeared in any training email labelled spam. What does unsmoothed Naive Bayes conclude about a spam email containing refund? The probability of refund given spam is zero, which zeroes out the entire spam score. Adding one to every count, called Laplace smoothing, fixes this.

4:2314. Word weights

Word weights — Text Classification from Start to Finish

A linear classifier gives each word a weight. Delicious pushes strongly towards positive, while slow and rude push towards negative. The weights add up, and a sigmoid turns the total into a probability. But look at the last sentence: not bad at all is positive, yet word by word it looks negative.

4:4515. Logistic regression

Logistic regression — Text Classification from Start to Finish

Logistic regression learns such weights by gradient descent, placing a boundary between the classes and assigning each point a probability. Here accuracy climbs from fifty nine to ninety three percent as training runs. On TF IDF features, it is one of the strongest simple baselines in text classification.

5:0616. A strong baseline

A strong baseline — Text Classification from Start to Finish

Here is a baseline that is hard to beat for its cost. A pipeline combines a TF IDF vectoriser, using words and bigrams, with logistic regression using balanced class weights. Five fold cross validation reports macro F1. It trains in seconds and gives a reference point for anything fancier.

5:2717. Cross-validation

Cross-validation — Text Classification from Start to Finish

Always evaluate with cross validation or a held out test set. With five folds, every example is used for validation exactly once. Here the scores range from point eight two to point eight seven, averaging point eight four eight: that spread tells you how much a single split could mislead.

5:4818. Reading the confusion matrix

Reading the confusion matrix — Text Classification from Start to Finish

The confusion matrix shows exactly which mistakes happen. For a spam filter tested on one hundred emails: forty two spam caught, eight real emails wrongly flagged, six spam missed and forty four real emails correctly passed, giving precision of eighty four percent and recall of eighty seven and a half.

6:0919. Imbalanced classes

Imbalanced classes — Text Classification from Start to Finish

Real label sets are often imbalanced. If ninety five percent of tickets are other, a model that always says other is ninety five percent accurate and completely useless. Use class weights or resampling, tune decision thresholds, and report per class precision and recall or macro F1.

6:2920. Pause and think

Pause and think — Text Classification from Start to Finish

Pause and think. A harmful content classifier flags posts for human review, and missing a harmful post is far worse than reviewing a harmless one. Which metric do you prioritise? Recall for the harmful class, while keeping precision acceptable so reviewers are not overwhelmed.

6:4821. Pre-trained transformers

Pre-trained transformers — Text Classification from Start to Finish

The next step up is a pre trained transformer such as BERT. It has already learned a great deal about language by predicting masked words from both sides of their context, so it understands negation, word order and synonyms far better than a bag of words.

7:0722. Fine-tuning

Fine-tuning — Text Classification from Start to Finish

We fine tune it for our labels: add a small classification head on top of the pre trained network and train on our labelled examples with a small learning rate. With a few thousand labelled texts, fine tuned transformers usually beat TF IDF baselines clearly, especially on subtle cases.

7:2823. Zero-shot and few-shot

Zero-shot and few-shot — Text Classification from Start to Finish

Large language models can classify with no training at all. Describe the labels in a prompt, zero shot, or add a handful of examples, few shot, and ask the model to choose. This is perfect for prototypes, brand new labels or when you have almost no labelled data.

7:4824. Zero-shot in code

Zero-shot in code — Text Classification from Start to Finish

Here is zero shot classification with an open model. A model trained on natural language inference judges whether the text entails each candidate label, such as this text is about billing. Give it the text and the candidate labels, and it returns a ranked list with scores, without any training data.

8:1025. Which approach?

Which approach? — Text Classification from Start to Finish

Which approach should you choose? A TF IDF linear model trains in seconds, runs anywhere and is transparent, but needs labels and misses subtlety. Fine tuned transformers understand context much better. Prompted language models need no labels at all, but cost more per prediction. Many teams prototype with prompting, then train a cheaper model.

8:3326. Error analysis

Error analysis — Text Classification from Start to Finish

The most valuable hour in any project is error analysis. Read a sample of misclassified examples and group them by cause: wrong labels, negation, sarcasm, missing domain vocabulary or genuinely ambiguous categories. Very often the best fix is in the data or the label definitions, not the model.

8:5327. Common failure modes

Common failure modes — Text Classification from Start to Finish

Common failures include negation and sarcasm, which fool bag of words models; domain shift, as new products and slang appear; label noise from inconsistent annotators; and spurious cues, such as an email signature that happens to predict the label. Each has a matching fix.

9:1228. Labelling smarter

Labelling smarter — Text Classification from Start to Finish

Labelling is expensive, so label smartly. With active learning, the current model chooses which unlabelled texts should be labelled next, usually those it is least certain about. Each new annotation then teaches the model as much as possible, often reaching the same accuracy with far fewer labels.

9:3229. Beyond English

Beyond English — Text Classification from Start to Finish

Many applications serve several languages. Multilingual transformers such as XLM R share one model across about a hundred languages, so a classifier trained with English labels can often work on Bangla or French too. Always check performance separately for each language, since quality varies.

9:5130. Deploy and monitor

Deploy and monitor — Text Classification from Start to Finish

After deployment, language keeps moving. New products, slang and events change the text your classifier sees. Monitor the mix of predicted labels and the model’s confidence, label a fresh sample regularly, and retrain. A sudden rise in low confidence predictions often signals a brand new topic.

10:1131. Best practices

Best practices — Text Classification from Start to Finish

In short: invest in clear labels, start with a TF IDF and logistic regression baseline, measure with macro F1 and per class metrics, read the errors before changing the model, and keep monitoring and refreshing, because language keeps changing.

10:2832. Recap

Recap — Text Classification from Start to Finish

To recap. Frame your labels carefully, because label quality caps model quality. TF IDF with n grams and a linear model gives a fast, strong baseline. Cross validate and track precision, recall and macro F1. Fine tuned transformers handle context, language models can classify zero shot, and error analysis plus monitoring keep a classifier useful.

Key takeaways

  • Text classification maps text to labels; clear labelling guidelines and annotator agreement set the ceiling on quality.
  • TF-IDF with word n-grams plus logistic regression (or Naive Bayes) is a fast, strong baseline.
  • Use cross-validation and per-class metrics (precision, recall, macro-F1); accuracy misleads on imbalanced labels.
  • Fine-tuned pre-trained transformers capture negation and context; LLMs can classify zero-/few-shot without training.
  • Error analysis usually reveals data and label problems: negation, sarcasm, domain shift, label noise, spurious cues.
  • Monitor deployed classifiers for drift in language and topics, and retrain with fresh labels.

Check yourself

  1. Why add bigram features such as “not good”?
    Show answer

    To capture short phrases and negation a bag of single words misses — “not good” differs from “good”.

  2. What does Naive Bayes assume?
    Show answer

    Words are independent given the class — A “naive” but effective independence assumption.

  3. With 95% of tickets labelled “other”, which metric is most informative?
    Show answer

    Macro-F1 / per-class recall and precision — Accuracy rewards always predicting “other”.

  4. What is zero-shot classification?
    Show answer

    Classifying using label descriptions without task-specific training examples — E.g. an NLI model or a prompted LLM.

  5. Error analysis mainly involves…
    Show answer

    Reading misclassified examples and grouping them by cause — It reveals what to fix next.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/text-classification-end-to-end.html