Course Hive
Search

Welcome

Sign in or create your account

Continue with Google
or
Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 2 - Transformer-Based Models & Tricks
Play lesson

Large Language Models (LLMs) - Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 2 - Transformer-Based Models & Tricks

4.0 (2)
25 learners

What you'll learn

Analyze case studies to identify core principles
Apply theoretical frameworks to real-world scenarios
Evaluate evidence to support arguments
Synthesize information from multiple sources

This course includes

  • 34.5 hours of video
  • Certificate of completion
  • Access on mobile and TV

Summary

Keywords

Full Transcript

For more information about Stanford’s graduate programs, visit: https://online.stanford.edu/graduate-education October 3, 2025 This lecture covers: • Attention approximation • MHA, MQA, GQA • Position embeddings (regular, learned) • RoPE and applications • Transformer-based architectures • BERT and its derivatives To follow along with the course schedule and syllabus, visit: https://cme295.stanford.edu/syllabus/ Chapters: 00:00:00 Introduction 00:01:30 Recap of Transformers 00:10:37 Overview of position embeddings 00:15:36 Sinusoidal embeddings 00:25:56 T5 bias, ALiBi 00:31:02 RoPE 00:43:42 Layer normalization 00:50:39 Sparse attention 00:55:38 Sharing attention heads 01:02:42 Transformer-based models 01:11:38 BERT deep dive 01:33:24 BERT finetuning 01:43:30 Extensions of BERT Afshine Amidi is an Adjunct Lecturer at Stanford University. Shervine Amidi is an Adjunct Lecturer at Stanford University. View the course playlist: https://www.youtube.com/playlist?list=PLoROMvodv4rOCXd21gf0CF4xr35yINeOy

Course Hive

Continue this lesson in the app

Install CourseHive on Android or iOS to keep learning while you move.

Related Courses

FAQs

Course Hive
Download CourseHive
Keep learning anywhere