Transformers & Self-Attention
🤚🏻 Welcome to Assignment 2 of the CV track of CSOC’26
Introduction
Modules are present on the website, complete one to get the access to the next one. Each module has its deadline, but you can do it before and move to the next. In order to get access to the next module beforehand, ask the mentors on discord to review your solution for the current module. If they think you’ve done it correctly, then you will be given the access to the next module.
Adhere strictly to deadlines. Submissions will be evaluated on approach, technical correctness, and clarity. The most technically accurate solution may not necessarily be the one chosen; clarity of thought and a well-reasoned approach will be valued more.
Communities
All resources and updates will be shared through the CSOC website and the Discord Community. Ensure you’re registered on the website and join the Discord community. In case we fail to respond on discord, you can contact us on Whatsapp:
- Sudarshan: 9967215737
- Vidit: 7398089536
- Abhyudaya: 8305750670
- Raghav: 9368833344
Resources
Let us begin by gaining an introductory understanding of Sequence to Sequence learning, Transformers and Self-Attention.
Dataset Generation
Sequence to Sequence Learning
- This video by statquest is a great start to know about seq to seq architecture.
Seq to seq explained.
- A video about encoder decoder, that dives deep into seq to seq : Encoder Decoder
- Check out these blog posts as well:
- Sequence to Sequence Learning with Neural Networks : Foundational paper introducing sequence-to-sequence learning.
Seq-to-Seq Learning
- CampusX Tutorial
Standard Transformers & Self-Attention
This section provides useful resources for learning the basics of Transformers:
- This playlist by Sir Mitesh Khapra explains all the componets of transformer very
well(Videos 1-16 are enough) - MUST WATCH
- The Original Paper: The seminal paper that started it all. Focus on understanding the Multi-Head Attention equations.
Attention Is All You Need (Vaswani et al.)
- The Illustrated Transformer**:** An intuitive, visual explanation of the Transformer architecture.
The Illustrated Transformer by Jay Alammar
- Andrej Karpathy’s “Let’s build GPT”: While focused on a decoder-only model for text generation, the first half provides explanation of self-attention mechanics from scratch.
Neural Networks: Zero to Hero
- CampusX: 100 Days of Deep Learning provides a deep insight into attention and transformer architecturs. Refer videos 67 ( 6-7) to 84
100 Days of Deep Learning
Deep Dive into Individual Components
If you are having a hard time understanding individual components of the Transformer, use these specific resources:
Hugging Face & Implementation
- Hugging Face Official Guide: Hugging face guide on transformers (For now, the introduction, models, and preprocessing sections will be enough. This teaches you how to use the library for NLP tasks).
Visualization Aid
Visualizing how attention heads route information can help the math click. Here are the best tools to see Transformers in action:
- 3Blue1Brown - Attention visually explained: Visual explanation of attention, dot products, and softmax. Attention iĂn neural networks
- LLM Visualization: Visualization of tensor flow and matrix multiplications across all Transformer layers. LLM Visualization
- Transformer Explainer: Interactive web tool to observe live token transformations layer by layer. Transformer Explainer
- BertViz: Reference repository for visualizing attention weights across layers and heads. BertViz Interactive Tool