Microsoft Models: Intelligence, Performance & Price
Artificial Analysis has benchmarked 4 models from Microsoft. Below is a comparison of the key metrics across these models.
- For intelligence, the top model from Microsoft is Phi-4 Mini Instruct at 6 (estimated).
- For output speed, the fastest model is Phi-4 Mini Instruct at 45 t/s.
- For latency, Phi-4 Multimodal Instruct at 0.79s offers the lowest time to first answer token.
Intelligence
Artificial Analysis Intelligence Index
Intelligence Index vs. Cost per Intelligence Index Task
Cost
Cost per Intelligence Index Task
Speed & Latency
Output Speed
Capability Scores
Capability Indexes
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AA-LCR v1.1, GDP.pdf, AutomationBench-AA · Higher is better
Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity's Last Exam, AutomationBench-AA · Higher is better
Incorporates 6 evaluations: AA-Omniscience, Humanity's Last Exam, CritPt, GDPval-AA v2.1, AA-Briefcase v1.1, Terminal-Bench 4.0 · Higher is better
Incorporates 5 evaluations: AA-Omniscience, Humanity's Last Exam, GDPval-AA v2.1, AA-Briefcase v1.1, AA-LCR v1.1 · Higher is better
All Microsoft Releases
Further details
Weights | Provider Benchmarks | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Phi-4 Mini Instruct | 6 | 3.8B | 128k | $0.0 | 45 | ||||
| Phi-4 | 6 | 14B | 16k | $0.2 | 40 | ||||
| Phi-3 Mini Instruct 3.8B | 6 | 3.8B | 4k | - | - | - | |||
| Phi-4 Multimodal Instruct | 6 | 5.6B | 128k | $0.0 | 17 |