What BERT Became: RoBERTa → DistilBERT → SBERT → ModernBERT
How the pretraining recipe evolved after 2018: more data and no NSP (RoBERTa), distillation, sentence embeddings, and the 2024 rewrite.
Playlist by @bert_pretraining_turn
8 tracks, shared on Audicious.
- RoBERTa model (BERT) in NLP explained — Data Science in your pocket
- RoBERTa: A Robustly Optimized BERT Pretraining Approach — Yannic Kilcher
- Knowledge Distillation in Deep Learning - DistilBERT Explained — Dingu Sagar
- Knowledge Distillation: How LLMs train each other — Julia Turc
- Sentence Transformers - EXPLAINED! — CodeEmporium
- SBERT (Sentence Transformers) is not BERT Sentence Embedding: Intro & Tutorial (#sbert Ep 37) — Discover AI
- 6 Years of AI Progress: ModernBERT Finally Replaces BERT — No Hype AI
- ModernBERT - Modern Replacement for BERT | RAG, Embeddings, Classification, Reranking — Venelin Valkov