Interview preparation / साक्षात्कार अभ्यास
Transformers
60 questions / प्रश्न · 8 sections
Transformer Fundamentals Q1–10
- What is Transformer architecture?Transformer architecture क्या है?
- Why is Transformer need padi important?Transformer की आवश्यकता क्यों padi?
- What is the difference between Transformer and RNN?Transformer और RNN में क्या अंतर है?
- What is the difference between Transformer and CNN?Transformer और CNN में क्या अंतर है?
- What are the main components of Transformer architecture?Transformer architecture के मुख्य components क्या हैं?
- What is Encoder and Decoder?Encoder और Decoder क्या होते हैं?
- What is the difference between Encoder-only, Decoder-only and Encoder-Decoder architecture?Encoder-only, Decoder-only और Encoder-Decoder architecture में क्या अंतर है?
- How would you approach Transformer sequence parallel process kiya jata?Transformer में sequence को parallel कैसे process किया jata है?
- Why is Transformer recurrence nahi important?Transformer में recurrence क्यों नहीं होती?
- Describe Transformer architecture ka complete data flow.Transformer architecture का complete data flow समझाएँ.
Attention Q11–20
- What is Attention mechanism?Attention mechanism क्या है?
- Why is Attention needed?Attention की आवश्यकता क्यों होती है?
- What is Query (Q), Key (K) and Value (V)?Query (Q), Key (K) और Value (V) क्या हैं?
- How would you approach Query, Key Value generate?Query, Key और Value कैसे generate होते हैं?
- What is Scaled Dot-Product Attention?Scaled Dot-Product Attention क्या है?
- What is Attention ka formula?Attention का formula क्या है?
- Explain QKᵀ represent.QKᵀ क्या represent करता है?
- Why is √dₖ divide important?√dₖ से divide क्यों करते हैं?
- Explain Softmax attention.Softmax attention में क्या करता है?
- Explain Attention weights represent.Attention weights क्या represent करते हैं?
Self-Attention Q21–28
- What is Self-Attention?Self-Attention क्या है?
- What is the difference between Self-Attention and normal Attention?Self-Attention और normal Attention में क्या अंतर है?
- Explain Self-Attention complete step-by-step process.Self-Attention का complete step-by-step process?
- Why is Q, K, V same input generate important?Q, K, V same input से क्यों generate होते हैं?
- How would you approach Self-Attention token apne baaki tokens attend?Self-Attention में token अपने baaki tokens को कैसे attend करता है?
- What is Self-Attention ka computational complexity?Self-Attention का computational complexity क्या है?
- Why is Long sequences Self-Attention expensive important?Long sequences के लिए Self-Attention expensive क्यों है?
- What is Self-Attention ki limitations?Self-Attention की limitations क्या हैं?
Multi-Head Attention Q29–35
- What is Multi-Head Attention?Multi-Head Attention क्या है?
- Why is Multi-Head Attention need important?Multi-Head Attention की आवश्यकता क्यों है?
- Explain Multiple attention heads learn.Multiple attention heads क्या learn करते हैं?
- How do Single-head and Multi-head attention differ?Single-head vs Multi-head attention?
- How would you approach Heads outputs combine?Heads के outputs को कैसे combine करते हैं?
- What is Multi-Head Attention ka computational cost?Multi-Head Attention का computational cost क्या है?
- Explain Agar number of heads increase karein effect.अगर number of heads increase karein तो क्या effect होगा?
Positional Encoding Q36–41
- Why is Transformer positional information important?Transformer को positional information की ज़रूरत क्यों है?
- What is Positional Encoding?Positional Encoding क्या है?
- What is Sinusoidal Positional Encoding?Sinusoidal Positional Encoding क्या है?
- Why is Sinusoidal functions kiye gaye important?Sinusoidal functions क्यों उपयोग kiye gaye?
- What is Learned Positional Embedding?Learned Positional Embedding क्या है?
- What is the difference between Positional Encoding and Positional Embedding?Positional Encoding और Positional Embedding में क्या अंतर है?
Transformer Encoder Q42–47
- What is Transformer Encoder?Transformer Encoder क्या है?
- What are the main components of Encoder block?Encoder block के मुख्य components क्या हैं?
- How would you approach Encoder Self-Attention?Encoder में Self-Attention कैसे उपयोग होता है?
- What is Feed-Forward Network (FFN)?Feed-Forward Network (FFN) क्या है?
- What is Residual Connection?Residual Connection क्या है?
- What is Layer Normalization ka role?Layer Normalization का role क्या है?
Transformer Decoder & Masking Q48–54
- What is Transformer Decoder?Transformer Decoder क्या है?
- What are the main components of Decoder block?Decoder block के मुख्य components क्या हैं?
- What is Encoder-Decoder Attention / Cross-Attention?Encoder-Decoder Attention / Cross-Attention क्या है?
- What is Masked Self-Attention?Masked Self-Attention क्या है?
- What is Causal Mask?Causal Mask क्या है?
- What is Padding Mask?Padding Mask क्या है?
- What is the difference between Causal Mask and Padding Mask?Causal Mask और Padding Mask में क्या अंतर है?
Training, Generation & Practical Q55–60
- How would you approach Transformer training?Transformer training कैसे होती है?
- What is Autoregressive generation?Autoregressive generation क्या है?
- How would you approach Transformer next token predict?Transformer next token कैसे predict करता है?
- How would you approach Transformer inference decoding?Transformer inference में decoding कैसे होती है?
- What is Transformer ki major limitations?Transformer की major limitations क्या हैं?
- Describe Complete Transformer architecture ko input se output tak.Complete Transformer architecture को input से output tak समझाएँ.