The Cloze Trick: Masked Language Modeling

The heart of the pretraining turn: hide 15% of the tokens (80/10/10), predict them bidirectionally, repeat on BookCorpus + Wikipedia.

Playlist by @bert_pretraining_turn

6 tracks, shared on Audicious.

  1. 🤗 Tasks: Masked Language Modeling — Hugging Face
  2. Masked Language Modeling (MLM) in BERT pretraining explained — Data Science in your pocket
  3. What is a Masked Language Model (MLM)? Masked vs. Causal AI — Eye on Tech
  4. BERT: Masked Language Modeling (Natural Language Processing at UT Austin) — Greg Durrett
  5. L21: Bert- masked language modelling & applications — IIT Madras - B.S. Degree Programme
  6. BERT Research - Ep. 8 - Inner Workings V - Masked Language Model — InnerWorkingsAI