The Cloze Trick: Masked Language Modeling
The heart of the pretraining turn: hide 15% of the tokens (80/10/10), predict them bidirectionally, repeat on BookCorpus + Wikipedia.
Playlist by @bert_pretraining_turn
6 tracks, shared on Audicious.
- 🤗 Tasks: Masked Language Modeling — Hugging Face
- Masked Language Modeling (MLM) in BERT pretraining explained — Data Science in your pocket
- What is a Masked Language Model (MLM)? Masked vs. Causal AI — Eye on Tech
- BERT: Masked Language Modeling (Natural Language Processing at UT Austin) — Greg Durrett
- L21: Bert- masked language modelling & applications — IIT Madras - B.S. Degree Programme
- BERT Research - Ep. 8 - Inner Workings V - Masked Language Model — InnerWorkingsAI