All MicroEvals
You are acting as a senior local-AI systems architect. Date...
Create MicroEval
Header image for You are acting as a senior local-AI systems architect.

Date...

You are acting as a senior local-AI systems architect. Date...

Prompt

You are acting as a senior local-AI systems architect. Date context: 27 August 2026. I need you to identify the strongest OPEN-WEIGHT / OPEN-SOURCE local LLMs for a persistent autonomous agent system running on this exact hardware: HARDWARE - Windows 11 - NVIDIA RTX 4060 8 GB VRAM - Ryzen 7 5700X - 64 GB system RAM - normally 30–40 GB RAM free because many other applications are running - local inference should preferably stay on GPU - CPU/RAM offload is allowed only if the capability gain is large enough to justify the latency - system may run for many hours continuously PROJECT I am building HYDRA / Factory Capo: a persistent autonomous AI assistant / agent system, not just a chatbot. The model will be used through an agent harness such as OpenClaw / OpenHands / equivalent and must be good at: 1. tool/function calling 2. strict JSON / structured output 3. multi-step agent workflows 4. ReAct-style tool use 5. planning before action 6. recovering after a failed tool call 7. reading and editing real repositories 8. coding and debugging 9. terminal interaction 10. long-running tasks 11. respecting tool permissions and sandbox boundaries 12. deciding when NOT to use a tool 13. choosing direct answer vs reasoning vs ReAct vs full workflow 14. low hallucinated-success rate 15. maintaining state across many turns 16. Italian + English 17. OpenAI-compatible API / LM Studio / llama.cpp / Ollama compatibility IMPORTANT: Do NOT rank models primarily by MMLU, GPQA, AIME or generic chat quality. This is an AGENTIC SYSTEM benchmark. The most important metrics are: - agent task success rate - tool selection accuracy - valid tool arguments - structured-output reliability - coding/repository capability - recovery after errors - unnecessary-tool avoidance - instruction following - fake-PASS / hallucinated-success rate - stability over multi-step loops - context reliability - real context feasible on RTX 4060 8 GB - tokens/sec - time-to-first-token - VRAM usage - RAM spill / CPU offload - GGUF quality at Q4_K_M / Q5_K_M - compatibility with agent harnesses - license Evaluate AT LEAST these candidates: - IBM Granite 4.2 8B - Ornith 1.5 9B - Qwen3.5 9B - Gemma 4 12B Also add any CURRENT 2026 open model that you believe is genuinely superior for this exact hardware. You may include larger models (12B–35B or MoE) only in a separate "RAM/OFFLOAD" category. Do NOT pretend that a model supporting 128K/262K/512K architecturally means that context is realistically usable on 8 GB VRAM. Estimate practical context for this machine separately. Do NOT assume native Chain-of-Thought automatically makes a model a better agent. Distinguish: MODEL REASONING CAPABILITY from HARNESS / ORCHESTRATION CAPABILITY. Do not reveal private chain-of-thought. Give concise decision rationale, evidence, assumptions and confidence instead. OUTPUT FORMAT A) TOP 10 LOCAL MODELS Table: Rank Model Parameters / architecture Best quantization for RTX 4060 8 GB Estimated VRAM Estimated additional RAM Practical context on this hardware Tool calling Structured JSON Coding Agent loops Error recovery Speed License Main weakness Overall HYDRA score /100 B) FULL-GPU WINNERS Give the best 3 models that can realistically operate mostly inside 8 GB VRAM. C) RAM-OFFLOAD WINNERS Give up to 3 larger models worth using with 30–40 GB system RAM. Clearly state expected latency penalty. D) ROLE-BASED ROUTING Choose the best local model for each: - everyday Factory Capo - ReAct/tool worker - coding worker - planner - structured-output worker - RAG/long-document worker - independent local validator A model may win multiple roles. E) DIRECT HEAD-TO-HEAD Compare: Granite 4.2 8B vs Ornith 1.5 9B vs Qwen3.5 9B vs Gemma 4 12B For each state: - where it clearly wins - where it loses - whether I should replace Ornith with it - confidence: HIGH / MEDIUM / LOW F) FINAL DECISION Return exactly: CURRENT_CHAMPION = CHALLENGER_1 = CHALLENGER_2 = BEST_AGENT_MODEL = BEST_CODE_MODEL = BEST_REASONING_MODEL = BEST_LONG_CONTEXT_MODEL = BEST_8GB_MODEL = BEST_RAM_OFFLOAD_MODEL = Then answer: "If this were your RTX 4060 8 GB machine running an autonomous agent 24/7, which model would you actually load first and why?" G) BENCHMARK PLAN Design a 20-task local benchmark to verify your recommendation experimentally. Include: - direct answer/no tool - one tool - multiple tools - forbidden tool - malformed JSON recovery - typed handoff - code edit - terminal failure - repair - RAG citation - long context - ambiguous goal - ask clarification - do not fake PASS - exactly-once action - bounded retry - multi-step planning - memory retrieval - resume - 30-minute agent workload The recommendation must ultimately be based on this benchmark, not reputation. If information about a new model is uncertain or you cannot verify it, explicitly write UNKNOWN instead of inventing specifications.

Drag to resize
Drag to resize
Drag to resize