
Benchmark mature, production-ready AI models for IntelFactor...
Prompt
Benchmark mature, production-ready AI models for IntelFactorβs core goal: **Customer order β factory evidence β bottleneck β intervention β measured result β economic impact.** Separate all results into four regions: * United States * European Union * Mainland China * Hong Kong Do not combine China and Hong Kong. Do **not** use the latest or newly released models. Only include models that have been generally available for at least 60 days, have stable production APIs, published pricing, and are not preview/beta/experimental. Freeze the model list for the benchmark. Test every model on the same IntelFactor-style tasks: * Extract facts from messy factory evidence * Preserve provenance and avoid hallucinations * Identify the true production bottleneck * Choose one evidence-backed intervention * Calculate baseline, throughput gain, cost savings, ROI, payback period, and economic impact * Reconcile conflicting factory data * Connect customer requirements to production constraints and customer impact * Complete multi-step workflows from evidence β analysis β decision β economic result Score each model on: * Evidence accuracy * Hallucination resistance * Bottleneck identification * Economic/math accuracy * Intervention quality * Manufacturing reasoning * Instruction following * Agentic task completion * Reliability * Cost per successful task * Time per successful task Use hard production gates: a model should not win if it hallucinates factory facts, makes material calculation errors, ignores conflicting evidence, or cannot complete the full workflow. Produce: 1. Best qualifying model from the US 2. Best qualifying model from the EU 3. Best qualifying model from mainland China 4. Best qualifying model from Hong Kong, or state if no eligible model exists 5. Overall best model for IntelFactor 6. Best quality/cost model 7. Recommended single-model or multi-model deployment strategy Prioritize **evidence β decision β intervention β measured economic result**, not generic benchmark intelligence.