All MicroEvals
LLM Comparison for 8D Process Support in Quality Management.
Create MicroEval
Header image for LLM Comparison for 8D Process Support in Quality Management.

LLM Comparison for 8D Process Support in Quality Management.

Research-based comparative analysis of Large Language Models (LLMs) for pre-selecting the 3 most suitable candidates for an AI-supported 8D -Agent in Quality Management.

Prompt

# Task Conduct a research-based comparative analysis of Large Language Models (LLMs) for pre-selecting the 3 most suitable candidates for an AI-supported 8D-Agent in Quality Management. # Context - Project: Development and evaluation of an AI-supported 8D-Agent for quality management in Advanced Semiconductor Industry. - Objective: Generate high-quality 8D reports from technical documentation through evidence-based interaction with Quality Engineers - Interaction Model: Based on Facts, Questions, and Trusted Sources - Final Selection: Top 3 models will undergo experimental evaluation with real quality cases. # Evaluation Criteria Assess each model against these criteria. Weightings must be supported by research sources (8D methodology standards, quality management literature, RCA frameworks) | Criterion | Evaluation Question | |-----------|--------|---------------------| | Root-Cause-Reasoning | Can the model logically derive technical root causes with traceable causal chains? | | Technical Plausibility | Does the model recognize realistic technical relationships in semiconductor/manufacturing contexts? | | Document Understanding | Can it process and extract from structured documents (PRF, QIR, BoM, test reports, technical specs)? | | Source Traceability | Does the model reliably cite sources and maintain reference accuracy? | | Question Quality | Can it generate relevant, prioritized clarifying questions when information is incomplete? | | Fact/Inference Distinction | Does it clearly separate confirmed facts from derived conclusions and flag uncertainties? | | Multimodality | Can it interpret technical drawings, images, screenshots, and diagrams? | | Causal Consistency | Does it maintain logical consistency across Problem → Root Cause → Actions? | | Actionability | Does it generate concrete, measurable, and implementable recommendations? | | Customer Orientation | Can it formulate content professionally and suitable for customer communication? | | Stability/Reproducibility | Does it produce consistent outputs with identical inputs? | | Efficiency | Response time and token efficiency for complex tasks | # Required Output Structure ## 1. Model Capability Matrix Create a comparison table rating each model per criterion (1-5 scale): - 5 = Excellent (documented strength, benchmark leader) - 4 = Good (above average, reliable) - 3 = Adequate (meets basic requirements) - 2 = Limited (notable weaknesses) - 1 = Poor (significant gaps) | Model | Root-Cause | Tech. Plausibility | Doc. Understanding | Source Tracing | Question Quality | Fact/Inference | Multimodal | Causal Consistency | Actionability | Customer Orient. | Stability | Efficiency | **Weighted Score** | |-------|------------|-------------------|-------------------|----------------|------------------|----------------|------------|-------------------|---------------|------------------|-----------|------------|-------------------| ## 2. Evidence-Based Justification per Model For each model, provide: - **Strengths** for 8D-Agent use case (with research/benchmark references) - **Weaknesses** or limitations relevant to the task - **Key differentiators** vs. other candidates - **Source references** (papers, benchmarks, technical reports, official documentation) ## 3. Criterion-Specific Analysis For the top 3 weighted criteria (Root-Cause-Reasoning, Technical Plausibility, Document Understanding, Source Traceability), provide: - Which models excel and why - Known limitations or failure modes - Relevant benchmark results or research findings ## 4. Multimodality Assessment Specifically evaluate vision/multimodal capabilities: - Which models support image input? - Quality of technical drawing interpretation - Ability to reference visual elements in text output ## 5. Risk Assessment For each model, identify: - Hallucination tendency (critical for fact-based 8D reports) - Context window limitations vs. typical document set size - Known issues with technical/domain-specific content ## 6. Final Recommendation Provide a ranked selection of the **Top 3 Models** with: - Weighted total score - Primary justification (2-3 sentences) - Specific suitability for 8D-Agent requirements - Any caveats or conditions for selection ## 7. Sources & References List all sources used for the assessment: - Academic papers - Official benchmarks (MMLU, HumanEval, etc.) - Technical documentation - Independent evaluations # Constraints - Base assessments on documented capabilities, benchmarks, and research – not assumptions - Clearly indicate when information is limited or unavailable for a model - Consider practical factors: API availability, cost, context window, response latency - Focus on capabilities relevant to the specific 8D-Agent use case, not general performance # Output Format Provide the complete analysis in a structured format suitable for inclusion in a Bachelor thesis