All MicroEvals
A serverless AI inference platform has 8 H200 GPUs. It serve...
Create MicroEval
Header image for A serverless AI inference platform has 8 H200 GPUs. It serve...

A serverless AI inference platform has 8 H200 GPUs. It serve...

Prompt

A serverless AI inference platform has 8 H200 GPUs. It serves three workloads: 1. Coding agents: bursty traffic, 20 requests at once, then idle for several minutes. 2. Customer-support chat: steady traffic, each conversation has up to 80,000 tokens of context. 3. Document analysis: occasional 200,000-token documents requiring accurate extraction. Design the best deployment strategy for these workloads. Explain: - How you would allocate GPU capacity - Which workloads should share infrastructure - How to handle bursts and idle periods - How to control latency and cost - The three biggest failure risks Keep the answer under 700 words. State any assumptions clearly and finish with a concise recommendation table.