Files
waggle-os/benchmarks/calibration/v6-kappa-recal/phase2-cold-probes.jsonl
Oleg Maslov 0c3e2ead3b
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled
moving
2026-09-02 10:10:29 +02:00

7 lines
4.9 KiB
JSON

{"instance_id": "locomo_conv-43_q155", "cell": "agentic", "provider": "minimax", "alias": "minimax-m27-via-openrouter", "routing": "openrouter_direct_http", "http_status": 200, "error": null, "latency_ms": 28396, "prompt_tokens": 480, "completion_tokens": 635, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer 'great' is a direct synonym for John's stated feeling ('It was great!') and captures the same positive sentiment as 'huge success' from the ground truth, making it an acceptable equivalent formulation.", "raw_text": "{\n \"verdict\": \"correct\",\n \"failure_mode\": null,\n \"rationale\": \"The model's answer 'great' is a direct synonym for John's stated feeling ('It was great!') and captures the same positive sentiment as 'huge success' from the ground truth, making it an acceptable equivalent formulation.\"\n}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-43_q155", "cell": "agentic", "provider": "kimi", "alias": "kimi-k26-direct", "routing": "moonshot_direct_http", "http_status": 0, "error": "TimeoutError: The read operation timed out", "latency_ms": 60112, "prompt_tokens": null, "completion_tokens": null, "parsed_verdict": null, "parsed_failure_mode": null, "parsed_rationale": null, "raw_text": "", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-50_q015", "cell": "agentic", "provider": "minimax", "alias": "minimax-m27-via-openrouter", "routing": "openrouter_direct_http", "http_status": 200, "error": null, "latency_ms": 17311, "prompt_tokens": 573, "completion_tokens": 558, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer captures the primary reason (relax) and adds a conceptually aligned detail (recharge), which is acceptable as minor elaboration that is not contradictory to the ground truth of relaxing and calming.", "raw_text": "{\"verdict\":\"correct\",\"failure_mode\":null,\"rationale\":\"The model's answer captures the primary reason (relax) and adds a conceptually aligned detail (recharge), which is acceptable as minor elaboration that is not contradictory to the ground truth of relaxing and calming.\"}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-50_q015", "cell": "agentic", "provider": "kimi", "alias": "kimi-k26-direct", "routing": "moonshot_direct_http", "http_status": 200, "error": null, "latency_ms": 15158, "prompt_tokens": 550, "completion_tokens": 351, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer captures the ground truth that Dave visits parks to relax, and 'recharge' is an acceptable synonymous formulation of the calming, restorative benefit described in the context without introducing any incorrect claims.", "raw_text": "{\n \"verdict\": \"correct\",\n \"failure_mode\": null,\n \"rationale\": \"The model's answer captures the ground truth that Dave visits parks to relax, and 'recharge' is an acceptable synonymous formulation of the calming, restorative benefit described in the context without introducing any incorrect claims.\"\n}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-43_q155", "cell": "full-context", "provider": "minimax", "alias": "minimax-m27-via-openrouter", "routing": "openrouter_direct_http", "http_status": 200, "error": null, "latency_ms": 25481, "prompt_tokens": 480, "completion_tokens": 518, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer 'great' is a direct synonym for 'a huge success' from the ground truth, both expressing John's positive feeling about jogging without pain, and contains no incorrect claims.", "raw_text": "{\n \"verdict\": \"correct\",\n \"failure_mode\": null,\n \"rationale\": \"The model's answer 'great' is a direct synonym for 'a huge success' from the ground truth, both expressing John's positive feeling about jogging without pain, and contains no incorrect claims.\"\n}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-43_q155", "cell": "full-context", "provider": "kimi", "alias": "kimi-k26-direct", "routing": "moonshot_direct_http", "http_status": 200, "error": null, "latency_ms": 47213, "prompt_tokens": 456, "completion_tokens": 1617, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer 'great' is explicitly stated in the ground-truth context as John's direct feeling about jogging without pain, making it factually correct and an acceptable equivalent formulation.", "raw_text": "{\n \"verdict\": \"correct\",\n \"failure_mode\": null,\n \"rationale\": \"The model's answer 'great' is explicitly stated in the ground-truth context as John's direct feeling about jogging without pain, making it factually correct and an acceptable equivalent formulation.\"\n}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}