moving
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled

This commit is contained in:
Oleg Maslov
2026-09-02 10:10:29 +02:00
commit 0c3e2ead3b
3841 changed files with 970576 additions and 0 deletions

View File

@@ -0,0 +1,6 @@
{"instance_id": "locomo_conv-43_q155", "cell": "agentic", "provider": "minimax", "alias": "minimax-m27-via-openrouter", "routing": "openrouter_direct_http", "http_status": 200, "error": null, "latency_ms": 28396, "prompt_tokens": 480, "completion_tokens": 635, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer 'great' is a direct synonym for John's stated feeling ('It was great!') and captures the same positive sentiment as 'huge success' from the ground truth, making it an acceptable equivalent formulation.", "raw_text": "{\n \"verdict\": \"correct\",\n \"failure_mode\": null,\n \"rationale\": \"The model's answer 'great' is a direct synonym for John's stated feeling ('It was great!') and captures the same positive sentiment as 'huge success' from the ground truth, making it an acceptable equivalent formulation.\"\n}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-43_q155", "cell": "agentic", "provider": "kimi", "alias": "kimi-k26-direct", "routing": "moonshot_direct_http", "http_status": 0, "error": "TimeoutError: The read operation timed out", "latency_ms": 60112, "prompt_tokens": null, "completion_tokens": null, "parsed_verdict": null, "parsed_failure_mode": null, "parsed_rationale": null, "raw_text": "", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-50_q015", "cell": "agentic", "provider": "minimax", "alias": "minimax-m27-via-openrouter", "routing": "openrouter_direct_http", "http_status": 200, "error": null, "latency_ms": 17311, "prompt_tokens": 573, "completion_tokens": 558, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer captures the primary reason (relax) and adds a conceptually aligned detail (recharge), which is acceptable as minor elaboration that is not contradictory to the ground truth of relaxing and calming.", "raw_text": "{\"verdict\":\"correct\",\"failure_mode\":null,\"rationale\":\"The model's answer captures the primary reason (relax) and adds a conceptually aligned detail (recharge), which is acceptable as minor elaboration that is not contradictory to the ground truth of relaxing and calming.\"}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-50_q015", "cell": "agentic", "provider": "kimi", "alias": "kimi-k26-direct", "routing": "moonshot_direct_http", "http_status": 200, "error": null, "latency_ms": 15158, "prompt_tokens": 550, "completion_tokens": 351, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer captures the ground truth that Dave visits parks to relax, and 'recharge' is an acceptable synonymous formulation of the calming, restorative benefit described in the context without introducing any incorrect claims.", "raw_text": "{\n \"verdict\": \"correct\",\n \"failure_mode\": null,\n \"rationale\": \"The model's answer captures the ground truth that Dave visits parks to relax, and 'recharge' is an acceptable synonymous formulation of the calming, restorative benefit described in the context without introducing any incorrect claims.\"\n}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-43_q155", "cell": "full-context", "provider": "minimax", "alias": "minimax-m27-via-openrouter", "routing": "openrouter_direct_http", "http_status": 200, "error": null, "latency_ms": 25481, "prompt_tokens": 480, "completion_tokens": 518, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer 'great' is a direct synonym for 'a huge success' from the ground truth, both expressing John's positive feeling about jogging without pain, and contains no incorrect claims.", "raw_text": "{\n \"verdict\": \"correct\",\n \"failure_mode\": null,\n \"rationale\": \"The model's answer 'great' is a direct synonym for 'a huge success' from the ground truth, both expressing John's positive feeling about jogging without pain, and contains no incorrect claims.\"\n}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}
{"instance_id": "locomo_conv-43_q155", "cell": "full-context", "provider": "kimi", "alias": "kimi-k26-direct", "routing": "moonshot_direct_http", "http_status": 200, "error": null, "latency_ms": 47213, "prompt_tokens": 456, "completion_tokens": 1617, "parsed_verdict": "correct", "parsed_failure_mode": null, "parsed_rationale": "The model's answer 'great' is explicitly stated in the ground-truth context as John's direct feeling about jogging without pain, making it factually correct and an acceptable equivalent formulation.", "raw_text": "{\n \"verdict\": \"correct\",\n \"failure_mode\": null,\n \"rationale\": \"The model's answer 'great' is explicitly stated in the ground-truth context as John's direct feeling about jogging without pain, making it factually correct and an acceptable equivalent formulation.\"\n}", "opus_verdict_ref": "correct", "gpt_verdict_ref": "incorrect"}