Summary
Keywords
Full Transcript
Let's do a deep dive into the Transformer Neural Network Architecture for language translation. ABOUT ME β Subscribe: https://www.youtube.com/c/CodeEmporium?sub_confirmation=1 π Medium Blog: https://medium.com/@dataemporium π» Github: https://github.com/ajhalthor π LinkedIn: https://www.linkedin.com/in/ajay-halthor-477974bb/ RESOURCES [ 1 π] Transformer Architecture Image :https://github.com/ajhalthor/Transformer-Neural-Network/blob/main/Transformer_Architecture_complete.png [2 π] draw.io version of the image for clarity: https://github.com/ajhalthor/Transformer-Neural-Network/blob/main/Transformer_Architecture_complete.drawio PLAYLISTS FROM MY CHANNEL β Transformers from scratch playlist: https://www.youtube.com/watch?v=QCJQG4DuHT0&list=PLTl9hO2Oobd97qfWC40gOSU8C0iu0m2l4 β ChatGPT Playlist of all other videos: https://youtube.com/playlist?list=PLTl9hO2Oobd9coYT6XsTraTBo4pL1j4HJ β Transformer Neural Networks: https://youtube.com/playlist?list=PLTl9hO2Oobd_bzXUpzKMKA3liq2kj6LfE β Convolutional Neural Networks: https://youtube.com/playlist?list=PLTl9hO2Oobd9U0XHz62Lw6EgIMkQpfz74 β The Math You Should Know : https://youtube.com/playlist?list=PLTl9hO2Oobd-_5sGLnbgE8Poer1Xjzz4h β Probability Theory for Machine Learning: https://youtube.com/playlist?list=PLTl9hO2Oobd9bPcq0fj91Jgk_-h1H_W3V β Coding Machine Learning: https://youtube.com/playlist?list=PLTl9hO2Oobd82vcsOnvCNzxrZOlrz3RiD MATH COURSES (7 day free trial) π Mathematics for Machine Learning: https://imp.i384100.net/MathML π Calculus: https://imp.i384100.net/Calculus π Statistics for Data Science: https://imp.i384100.net/AdvancedStatistics π Bayesian Statistics: https://imp.i384100.net/BayesianStatistics π Linear Algebra: https://imp.i384100.net/LinearAlgebra π Probability: https://imp.i384100.net/Probability OTHER RELATED COURSES (7 day free trial) π β Deep Learning Specialization: https://imp.i384100.net/Deep-Learning π Python for Everybody: https://imp.i384100.net/python π MLOps Course: https://imp.i384100.net/MLOps π Natural Language Processing (NLP): https://imp.i384100.net/NLP π Machine Learning in Production: https://imp.i384100.net/MLProduction π Data Science Specialization: https://imp.i384100.net/DataScience π Tensorflow: https://imp.i384100.net/Tensorflow TIMESTAMPS 0:00 Introduction 1:38 Transformer at a high level 4:15 Why Batch Data? Why Fixed Length Sequence? 6:13 Embeddings 7:00 Positional Encodings 7:58 Query, Key and Value vectors 9:19 Masked Multi Head Self Attention 14:46 Residual Connections 15:50 Layer Normalization 17:57 Decoder 20:12 Masked Multi Head Cross Attention 22:47 24:03 Tokenization & Generating the next translated word 26:00 Transformer Inference Example
