What BERT Became: RoBERTa → DistilBERT → SBERT → ModernBERT

How the pretraining recipe evolved after 2018: more data and no NSP (RoBERTa), distillation, sentence embeddings, and the 2024 rewrite.

Playlist by @bert_pretraining_turn

8 tracks, shared on Audicious.

  1. RoBERTa model (BERT) in NLP explained — Data Science in your pocket
  2. RoBERTa: A Robustly Optimized BERT Pretraining Approach — Yannic Kilcher
  3. Knowledge Distillation in Deep Learning - DistilBERT Explained — Dingu Sagar
  4. Knowledge Distillation: How LLMs train each other — Julia Turc
  5. Sentence Transformers - EXPLAINED! — CodeEmporium
  6. SBERT (Sentence Transformers) is not BERT Sentence Embedding: Intro & Tutorial (#sbert Ep 37) — Discover AI
  7. 6 Years of AI Progress: ModernBERT Finally Replaces BERT — No Hype AI
  8. ModernBERT - Modern Replacement for BERT | RAG, Embeddings, Classification, Reranking — Venelin Valkov