
LLM Comparison for 8D Process Support in Quality Management.
Research-based comparative analysis of Large Language Models (LLMs) for pre-selecting the 3 most suitable candidates for an AI-supported 8D -Agent in Quality Management.
Prompt
# Task Conduct a research-based comparative analysis of Large Language Models (LLMs) for pre-selecting the 3 most suitable candidates for an AI-supported 8D-Agent in Quality Management. # Context - Project: Development and evaluation of an AI-supported 8D-Agent for quality management in Advanced Semiconductor Industry. - Objective: Generate high-quality 8D reports from technical documentation through evidence-based interaction with Quality Engineers - Interaction Model: Based on Facts, Questions, and Trusted Sources - Final Selection: Top 3 models will undergo experimental evaluation with real quality cases. # Evaluation Criteria Assess each model against these criteria. Weightings must be supported by research sources (8D methodology standards, quality management literature, RCA frameworks) | Criterion | Evaluation Question | |-----------|--------|---------------------| | Root-Cause-Reasoning | Can the model logically derive technical root causes with traceable causal chains? | | Technical Plausibility | Does the model recognize realistic technical relationships in semiconductor/manufacturing contexts? | | Document Understanding | Can it process and extract from structured documents (PRF, QIR, BoM, test reports, technical specs)? | | Source Traceability | Does the model reliably cite sources and maintain reference accuracy? | | Question Quality | Can it generate relevant, prioritized clarifying questions when information is incomplete? | | Fact/Inference Distinction | Does it clearly separate confirmed facts from derived conclusions and flag uncertainties? | | Multimodality | Can it interpret technical drawings, images, screenshots, and diagrams? | | Causal Consistency | Does it maintain logical consistency across Problem → Root Cause → Actions? | | Actionability | Does it generate concrete, measurable, and implementable recommendations? | | Customer Orientation | Can it formulate content professionally and suitable for customer communication? | | Stability/Reproducibility | Does it produce consistent outputs with identical inputs? | | Efficiency | Response time and token efficiency for complex tasks | # Required Output Structure ## 1. Model Capability Matrix Create a comparison table rating each model per criterion (1-5 scale): - 5 = Excellent (documented strength, benchmark leader) - 4 = Good (above average, reliable) - 3 = Adequate (meets basic requirements) - 2 = Limited (notable weaknesses) - 1 = Poor (significant gaps) | Model | Root-Cause | Tech. Plausibility | Doc. Understanding | Source Tracing | Question Quality | Fact/Inference | Multimodal | Causal Consistency | Actionability | Customer Orient. | Stability | Efficiency | **Weighted Score** | |-------|------------|-------------------|-------------------|----------------|------------------|----------------|------------|-------------------|---------------|------------------|-----------|------------|-------------------| ## 2. Evidence-Based Justification per Model For each model, provide: - **Strengths** for 8D-Agent use case (with research/benchmark references) - **Weaknesses** or limitations relevant to the task - **Key differentiators** vs. other candidates - **Source references** (papers, benchmarks, technical reports, official documentation) ## 3. Criterion-Specific Analysis For the top 3 weighted criteria (Root-Cause-Reasoning, Technical Plausibility, Document Understanding, Source Traceability), provide: - Which models excel and why - Known limitations or failure modes - Relevant benchmark results or research findings ## 4. Multimodality Assessment Specifically evaluate vision/multimodal capabilities: - Which models support image input? - Quality of technical drawing interpretation - Ability to reference visual elements in text output ## 5. Risk Assessment For each model, identify: - Hallucination tendency (critical for fact-based 8D reports) - Context window limitations vs. typical document set size - Known issues with technical/domain-specific content ## 6. Final Recommendation Provide a ranked selection of the **Top 3 Models** with: - Weighted total score - Primary justification (2-3 sentences) - Specific suitability for 8D-Agent requirements - Any caveats or conditions for selection ## 7. Sources & References List all sources used for the assessment: - Academic papers - Official benchmarks (MMLU, HumanEval, etc.) - Technical documentation - Independent evaluations # Constraints - Base assessments on documented capabilities, benchmarks, and research – not assumptions - Clearly indicate when information is limited or unavailable for a model - Consider practical factors: API availability, cost, context window, response latency - Focus on capabilities relevant to the specific 8D-Agent use case, not general performance # Output Format Provide the complete analysis in a structured format suitable for inclusion in a Bachelor thesis