← All topics / सभी विषय
Interview preparation / साक्षात्कार अभ्यास

Transformers

60 questions / प्रश्न · 8 sections

Transformer Fundamentals Q1–10

  1. What is Transformer architecture?Transformer architecture क्या है?
  2. Why is Transformer need padi important?Transformer की आवश्यकता क्यों padi?
  3. What is the difference between Transformer and RNN?Transformer और RNN में क्या अंतर है?
  4. What is the difference between Transformer and CNN?Transformer और CNN में क्या अंतर है?
  5. What are the main components of Transformer architecture?Transformer architecture के मुख्य components क्या हैं?
  6. What is Encoder and Decoder?Encoder और Decoder क्या होते हैं?
  7. What is the difference between Encoder-only, Decoder-only and Encoder-Decoder architecture?Encoder-only, Decoder-only और Encoder-Decoder architecture में क्या अंतर है?
  8. How would you approach Transformer sequence parallel process kiya jata?Transformer में sequence को parallel कैसे process किया jata है?
  9. Why is Transformer recurrence nahi important?Transformer में recurrence क्यों नहीं होती?
  10. Describe Transformer architecture ka complete data flow.Transformer architecture का complete data flow समझाएँ.

Attention Q11–20

  1. What is Attention mechanism?Attention mechanism क्या है?
  2. Why is Attention needed?Attention की आवश्यकता क्यों होती है?
  3. What is Query (Q), Key (K) and Value (V)?Query (Q), Key (K) और Value (V) क्या हैं?
  4. How would you approach Query, Key Value generate?Query, Key और Value कैसे generate होते हैं?
  5. What is Scaled Dot-Product Attention?Scaled Dot-Product Attention क्या है?
  6. What is Attention ka formula?Attention का formula क्या है?
  7. Explain QKᵀ represent.QKᵀ क्या represent करता है?
  8. Why is √dₖ divide important?√dₖ से divide क्यों करते हैं?
  9. Explain Softmax attention.Softmax attention में क्या करता है?
  10. Explain Attention weights represent.Attention weights क्या represent करते हैं?

Self-Attention Q21–28

  1. What is Self-Attention?Self-Attention क्या है?
  2. What is the difference between Self-Attention and normal Attention?Self-Attention और normal Attention में क्या अंतर है?
  3. Explain Self-Attention complete step-by-step process.Self-Attention का complete step-by-step process?
  4. Why is Q, K, V same input generate important?Q, K, V same input से क्यों generate होते हैं?
  5. How would you approach Self-Attention token apne baaki tokens attend?Self-Attention में token अपने baaki tokens को कैसे attend करता है?
  6. What is Self-Attention ka computational complexity?Self-Attention का computational complexity क्या है?
  7. Why is Long sequences Self-Attention expensive important?Long sequences के लिए Self-Attention expensive क्यों है?
  8. What is Self-Attention ki limitations?Self-Attention की limitations क्या हैं?

Multi-Head Attention Q29–35

  1. What is Multi-Head Attention?Multi-Head Attention क्या है?
  2. Why is Multi-Head Attention need important?Multi-Head Attention की आवश्यकता क्यों है?
  3. Explain Multiple attention heads learn.Multiple attention heads क्या learn करते हैं?
  4. How do Single-head and Multi-head attention differ?Single-head vs Multi-head attention?
  5. How would you approach Heads outputs combine?Heads के outputs को कैसे combine करते हैं?
  6. What is Multi-Head Attention ka computational cost?Multi-Head Attention का computational cost क्या है?
  7. Explain Agar number of heads increase karein effect.अगर number of heads increase karein तो क्या effect होगा?

Positional Encoding Q36–41

  1. Why is Transformer positional information important?Transformer को positional information की ज़रूरत क्यों है?
  2. What is Positional Encoding?Positional Encoding क्या है?
  3. What is Sinusoidal Positional Encoding?Sinusoidal Positional Encoding क्या है?
  4. Why is Sinusoidal functions kiye gaye important?Sinusoidal functions क्यों उपयोग kiye gaye?
  5. What is Learned Positional Embedding?Learned Positional Embedding क्या है?
  6. What is the difference between Positional Encoding and Positional Embedding?Positional Encoding और Positional Embedding में क्या अंतर है?

Transformer Encoder Q42–47

  1. What is Transformer Encoder?Transformer Encoder क्या है?
  2. What are the main components of Encoder block?Encoder block के मुख्य components क्या हैं?
  3. How would you approach Encoder Self-Attention?Encoder में Self-Attention कैसे उपयोग होता है?
  4. What is Feed-Forward Network (FFN)?Feed-Forward Network (FFN) क्या है?
  5. What is Residual Connection?Residual Connection क्या है?
  6. What is Layer Normalization ka role?Layer Normalization का role क्या है?

Transformer Decoder & Masking Q48–54

  1. What is Transformer Decoder?Transformer Decoder क्या है?
  2. What are the main components of Decoder block?Decoder block के मुख्य components क्या हैं?
  3. What is Encoder-Decoder Attention / Cross-Attention?Encoder-Decoder Attention / Cross-Attention क्या है?
  4. What is Masked Self-Attention?Masked Self-Attention क्या है?
  5. What is Causal Mask?Causal Mask क्या है?
  6. What is Padding Mask?Padding Mask क्या है?
  7. What is the difference between Causal Mask and Padding Mask?Causal Mask और Padding Mask में क्या अंतर है?

Training, Generation & Practical Q55–60

  1. How would you approach Transformer training?Transformer training कैसे होती है?
  2. What is Autoregressive generation?Autoregressive generation क्या है?
  3. How would you approach Transformer next token predict?Transformer next token कैसे predict करता है?
  4. How would you approach Transformer inference decoding?Transformer inference में decoding कैसे होती है?
  5. What is Transformer ki major limitations?Transformer की major limitations क्या हैं?
  6. Describe Complete Transformer architecture ko input se output tak.Complete Transformer architecture को input से output tak समझाएँ.