When moving the RAG system to production workloads, what specific bottlenecks did you encounter in the vector indexing or retrieval phase, and how did you optimize the latency for real-time conversational use ?
Indexing Bottle Neck
Retrival Bottle Neck
Generation Bottle Neck
