AI in Motion

Named Entity Recognition

Natural Language ProcessingBeginner1:115 chapters

Find people, organisations, places and dates in text and tag every token with BIO labels.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 In BIO tagging, what does “I-LOC” mean?
Q2 Why is “Apple” hard for NER?
Q3 Which tag is given to tokens that are not part of any entity?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. News articles, contracts and medical notes are full of names, places and dates. Named entity recognition finds them automatically.

Tagging entities. Here the model reads a sentence and labels each token. Ada Lovelace is a person, Charles Babbage is another person, London is a location and 1843 is a date. Each token gets a BIO tag. B marks the beginning of an entity, I marks its continuation, and O means outside any entity.

A sequence problem. NER is a sequence labelling problem. A token’s label depends on its neighbours: New York is one place. The same word can be a company or a fruit. So models read the whole sentence, from classic conditional random fields to modern fine-tuned Transformers.

Applications. NER helps search engines understand queries, extracts parties and amounts from legal and financial documents, finds drugs and symptoms in medical notes, and builds knowledge graphs.

Recap. To recap. NER labels entities in text using BIO tags. Context decides ambiguous cases, and fine-tuned Transformers are the modern approach.