Golden set v1
ชุดทดสอบหลัก 12 ทักษะ
โมเดลไหนทำงาน AI หลักทั้ง 12 ทักษะของ NoteStorming ได้ดีพอ ในราคาที่ถูกที่สุด
- Run
golden-v1-2026-08-31-001- ผู้ตัดสิน
anthropic/claude-sonnet-4.6- ขนาด
- 111 เคส × 3 รอบ
- รายงาน
benchmark/benchmark-report.md
| โมเดล | คุณภาพ | Critical | Hallucination | P50 | ต้นทุน / 1K | §31 score |
|---|---|---|---|---|---|---|
| Gemini 2.5 Flash | 89.2% | 92.8% | 2.1% | 1.0s | $0.34 | 0.925 |
| GPT-5.6 Luna | 90.2% | 94.2% | 1.8% | 2.4s | $0.23 | 0.919 |
| Gemini 3.7 Flash | 91.7% | 95.9% | 1.2% | 3.8s | $1.74 | 0.879 |
| GLM 5.3 Flash | 90.9% | 93.1% | 4.2% | 7.8s | $0.25 | 0.878 |
| DeepSeek V4 Flash | 88.3% | 95.9% | 4.2% | 4.2s | $0.07 | 0.872 |
| Qwen 3.7 Flash | 80.6% | 83.4% | 10.5% | 6.7s | $0.18 | 0.822 |
| GLM 5.2 | 68.0% | 70.8% | 25.5% | 2.7s | $1.80 | 0.736 |
| Kimi K3 | 90.9% | 93.5% | 4.5% | 12.5s | $9.21 | 0.631 |
ผู้ชนะตามเกณฑ์: Gemini 2.5 Flash
คะแนนรวมสูงสุดในตารางผลรวม (§31 score 0.925)
ระบบจริงไม่ได้ใช้โมเดลเดียว แต่ละทักษะใช้โมเดลที่ถูกที่สุดที่ผ่านเกณฑ์เข้ม (คุณภาพ ≥90%, critical ≥95%, hallucination และ schema fail ≤1%) ซึ่งมีโมเดลผ่านเกณฑ์นี้ 9 จาก 12 ทักษะ
Talk to Memory
คุยกับความจำของ Workspace
โมเดลไหนตอบคำถามจากความจำของ Workspace ได้ถูกต้อง ไม่แต่งข้อมูลขึ้นเอง และบอกตรง ๆ เมื่อเรื่องนั้นไม่อยู่ในความจำ
- Run
memory-chat-2026-10-08-001- ผู้ตัดสิน
google/gemini-2.5-pro- ขนาด
- 24 เคส × 3 รอบ
- รายงาน
benchmark/memory-chat/runs/memory-chat-2026-10-08-001/report.md
| โมเดล | คุณภาพ | Critical | Hallucination | P50 | ต้นทุน / 1K | Refusal (n/N) |
|---|---|---|---|---|---|---|
| GPT-5.6 Luna | 96.5% | 96.3% | 0.0% | 2.6s | $0.89 | 100.0% (15/15) |
| Claude Haiku 5.5 | 92.2% | 85.2% | 5.6% | 3.3s | $1.00 | 100.0% (15/15) |
| DeepSeek V4 Flash | 93.4% | 81.5% | 0.0% | 8.9s | $0.69 | 100.0% (15/15) |
ผู้ชนะตามเกณฑ์: DeepSeek V4 Flash
ถูกที่สุดในกลุ่มที่ผ่านทุกเกณฑ์ ได้แก่ คุณภาพ ≥90%, hallucination ≤2% และปฏิเสธคำถามที่ไม่มีในความจำได้ 100%
ระบบจริงใช้ GPT-5.6 Luna แทน เพราะมีคนรอคำตอบอยู่ (P50 2.6s เทียบกับ 8.9s, critical pass 96.3% เทียบกับ 81.5%) แลกกับค่าใช้จ่ายเพิ่ม $0.20 ต่อ 1K คำตอบ
Terminology context
คำศัพท์เฉพาะของทีม
คำศัพท์ที่ทีมกำหนดไปถึงทุก prompt หรือไม่ และ AI ใช้ตัวสะกดหลักตามที่ทีมกำหนดในความรู้ บทสรุป และคำตอบหรือเปล่า
- Run
terminology-2026-10-09-001- ผู้ตัดสิน
- ไม่มี ตรวจด้วยการเทียบสตริง
- ขนาด
- 20 ข้อตรวจ × 3 รอบ
- รายงาน
benchmark/terminology/runs/terminology-2026-10-09-001/report.md
- checks
- 20
- trials
- 3
- pass
- 60/60
Session diagram
ไดอะแกรมของ Session
โมเดลไหนวาดไดอะแกรมขั้นตอน ลำดับ และวงจรสถานะจากบทสนทนาได้ครบถ้วนที่สุด โดยไม่แต่งขั้นตอนที่ไม่มีใครพูดถึง
- Run
diagram-2026-10-09-001- ผู้ตัดสิน
- ไม่มี ตรวจด้วยการเทียบสตริง
- ขนาด
- 24 เคส × 3 รอบ
- รายงาน
benchmark/diagram/runs/diagram-2026-10-09-001/report.md
| โมเดล | P50 | ต้นทุน / 1K | Node recall | Invented (drawable) | None pass (n/N) |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | 11.3s | $6.80 | 84.7% | 1.8% | 58.3% (7/12) |
| Gemini 2.5 Flash | 5.6s | $2.70 | 83.7% | 0.8% | 0.0% (0/12) |
| GPT-5.6 Luna | 14.5s | $2.10 | 90.6% | 2.5% | 0.0% (0/12) |
| DeepSeek V4 Flash | 38.9s | $6.07 | 79.3% | 2.3% | 16.7% (2/12) |
ผู้ชนะตามเกณฑ์: GPT-5.6 Luna
node recall สูงสุดในกลุ่มที่ invented ≤5% บนบทสนทนาที่วาดได้ และ compile ไม่ล้มเหลว; เสมอกันตัดสินที่ราคา
คะแนนให้โดยการจับคู่โหนดและเส้นเชื่อมกับเฉลยแบบกำหนดได้ ไม่มีโมเดลตัดสิน; Luna ยังวาดเมื่อไม่มีอะไรให้วาด (0/12) ซึ่ง Gemini 3.7 Flash ทำได้ดีกว่า (7/12)
โมเดลที่ใช้จริงในแต่ละงาน
ค่าจาก apps/api/config/ai-router.json ที่ระบบใช้ส่งงาน AI แต่ละประเภทไปยังโมเดล ถ้าคำตอบของโมเดลหลักไม่ผ่านการตรวจรูปแบบ ระบบจะลองใหม่กับโมเดลสำรอง ซึ่งค่าตั้งต้นคือ google/gemini-3.7-flash
| งาน / โมเดล | เหตุผล |
|---|---|
batch_processgoogle/gemini-2.5-flash | Realtime batched classify+extract+summarize. Constituent-skill winners disagree (GLM-5.3-flash / GPT-5.6-Luna / none), so the realtime path… |
focus_analyzegoogle/gemini-3.7-flash | Strict benchmark winner: 95.9% critical quality, 1.2% hallucination — the lowest in the pool. |
session_summarygoogle/gemini-3.7-flash | No model passed the summarize bar strictly; end-of-session summaries are latency-insensitive and quality-critical, so they run on the… |
understand_templatemoonshotai/kimi-k3 | Per-skill benchmark winner. Rare, async operation (once per template upload), so its cost/latency profile is acceptable. |
prepare_documentdeepseek/deepseek-v4-flash | Strict winner at $0.07/1K with 95.9% critical quality. Async operation; the escalation chain covers its lower run-to-run consistency. |
session_diagramopenai/gpt-5.6-luna | Live session diagram board. Benchmark diagram-2026-10-09-001 (24 cases x 3 runs x 4 models, deterministic node/edge scoring, prompt… |
memory_chatopenai/gpt-5.6-luna | Talk to Memory (workspace assistant). Benchmark memory-chat-2026-10-08-001 (24 cases x 3 runs, judge gemini-2.5-pro, prompt imported from… |
transcribegoogle/gemini-2.5-flash | Composer voice input (speech to editable text). Audio-capable, Thai-capable, cheapest audio price in the catalog ($1 per million audio… |