← All topics / सभी विषय
Interview preparation / साक्षात्कार अभ्यास

LLMOps

60 questions / प्रश्न · 6 sections

Fundamentals Q1–10

  1. What is LLMOps?LLMOps क्या है?
  2. What is the difference between MLOps and LLMOps?MLOps और LLMOps में क्या अंतर है?
  3. What is LLM application ko production mein deploy karne ke major steps?LLM application को production में deploy करने के major steps क्या हैं?
  4. What is the difference between Development, staging and production environment?Development, staging और production environment में क्या अंतर है?
  5. Explain LLM application production components kaun.LLM application के production components कौन से होते हैं?
  6. What is Model serving?Model serving क्या है?
  7. What is Inference server?Inference server क्या होता है?
  8. What is LLM API service?LLM API service क्या है?
  9. How do Synchronous and asynchronous inference differ?Synchronous vs asynchronous inference?
  10. Describe Production LLM architecture.Production LLM architecture समझाएँ.

Model Serving & Inference Q11–20

  1. What is LLM inference?LLM inference क्या है?
  2. How do Self-hosted model and API-based model differ?Self-hosted model vs API-based model?
  3. How do you deploy Local LLM ko production mein?Local LLM को production में कैसे deploy करेंगे?
  4. Explain GPU inference kab.GPU inference की ज़रूरत कब होती है?
  5. How do CPU and GPU inference differ?CPU vs GPU inference?
  6. How do you optimize Model loading?Model loading कैसे optimize करेंगे?
  7. Why is Quantization production mein important?Quantization production में क्यों useful है?
  8. What is Batching?Batching क्या है?
  9. What is Continuous batching?Continuous batching क्या है?
  10. What is Streaming inference?Streaming inference क्या है?

Performance & Scaling Q21–30

  1. How do you reduce LLM latency ko?LLM latency को कैसे reduce करेंगे?
  2. What is TTFT?TTFT क्या है?
  3. What is Tokens/sec?Tokens/sec क्या है?
  4. What is Throughput?Throughput क्या है?
  5. How do you handle Concurrent requests?Concurrent requests कैसे handle करेंगे?
  6. How do Horizontal and vertical scaling differ?Horizontal vs vertical scaling?
  7. What is Load balancing?Load balancing क्या है?
  8. Why is Request queue important?Request queue क्यों उपयोग करते हैं?
  9. How would you approach Caching LLM applications?Caching LLM applications में कैसे उपयोग होती है?
  10. How would you approach High-traffic LLM API scale?High-traffic LLM API को scale कैसे करेंगे?

Cost & Reliability Q31–40

  1. Explain LLM application cost kin factors par depend karti.LLM application की cost kin factors पर depend करती है?
  2. How do you optimize Token cost?Token cost कैसे optimize करेंगे?
  3. How do Small and large model routing differ?Small vs large model routing?
  4. What is Model fallback?Model fallback क्या है?
  5. What is Retry strategy?Retry strategy क्या है?
  6. What is Exponential backoff?Exponential backoff क्या है?
  7. What is Rate limiting?Rate limiting क्या है?
  8. What is Circuit breaker?Circuit breaker क्या है?
  9. How would you approach Timeout handling?Timeout handling कैसे करेंगे?
  10. Describe Production LLM failure recovery architecture.Production LLM failure recovery architecture समझाएँ.

Evaluation & Monitoring Q41–50

  1. How would you approach Production LLM application monitor?Production LLM application को monitor कैसे करेंगे?
  2. What is Quality monitoring?Quality monitoring क्या है?
  3. Latency monitoring?Latency monitoring?
  4. Token/cost monitoring?Token/cost monitoring?
  5. Error-rate monitoring?Error-rate monitoring?
  6. What is Continuous evaluation?Continuous evaluation क्या है?
  7. What is Regression testing?Regression testing क्या है?
  8. How do you use User feedback ko production improvement mein?User feedback को production improvement में कैसे उपयोग करेंगे?
  9. Explain LangSmith LLMOps role.LangSmith का LLMOps में क्या role है?
  10. Explain Production AI system kaun- metrics track.Production AI system के लिए कौन से metrics track करेंगे?

Security & System Design Q51–60

  1. How would you approach Production LLM application secure?Production LLM application को secure कैसे करेंगे?
  2. How do API authentication and authorization differ?API authentication vs authorization?
  3. How do you handle Prompt injection ko production mein?Prompt injection को production में कैसे handle करेंगे?
  4. How would you approach PII protection?PII protection कैसे करेंगे?
  5. How do you manage Secrets/API keys ko securely?Secrets/API keys को securely कैसे manage करेंगे?
  6. How do you design Multi-user LLM application ka architecture?Multi-user LLM application का architecture कैसे डिज़ाइन करेंगे?
  7. How do you design Multi-tenant AI application kya hai aur?Multi-tenant AI application क्या है और कैसे डिज़ाइन करेंगे?
  8. Describe RAG + LLM + API ka production architecture.RAG + LLM + API का production architecture डिज़ाइन करें.
  9. Describe Agent + RAG + Tools + LLMOps ka architecture.Agent + RAG + Tools + LLMOps का architecture डिज़ाइन करें.
  10. Ek complete production-grade GenAI system design—from user request to monitoring.एक complete production-grade GenAI system डिज़ाइन करें—from user request तो monitoring.