6 lines
9.6 KiB
JSON
6 lines
9.6 KiB
JSON
{"request": {"contents": [{"role": "user", "parts": [{"text": "You are evaluating whether an LLM's answer is correct against ground truth.\n\n## Question\nWhen did Caroline go to the LGBTQ support group?\n\n## Ground-truth answer\n7 May 2023\n\n## Ground-truth supporting context (excerpt shown to the model)\nSession 1 (1:56 pm on 8 May, 2023):\nCaroline: I went to a LGBTQ support group yesterday and it was so powerful.\n\n## Model's answer\nunknown\n\n## Your task\n\nStep 1: Determine if the model's answer is correct.\n- \"correct\" means the model's answer contains all required facts from ground truth, with no additional incorrect claims.\n- Minor phrasing differences, synonyms, or alternative but equivalent formulations are acceptable.\n- Extra detail is acceptable ONLY if it is factually correct.\n\nStep 2: If incorrect, assign exactly one failure mode using this decision tree:\n\n1. Does the model explicitly refuse or say it does not know? ->F1 (ABSTAIN)\n2. Does the model answer a DIFFERENT question than was asked (coherent but off-topic)? ->F5 (OFF-TOPIC)\n3. Does the model rely on entities, names, dates, or claims that do NOT appear in the ground-truth context (fabrication)? ->F4 (HALLUCINATED)\n4. Does the model correctly state SOME required facts but miss others, without stating any incorrect facts? ->F2 (PARTIAL)\n5. Otherwise (model states facts derived from the context but gets them wrong): ->F3 (INCORRECT)\n\nStep 3: Return JSON only, no prose, in this exact schema:\n\n{\n \"verdict\": \"correct\" | \"incorrect\",\n \"failure_mode\": null | \"F1\" | \"F2\" | \"F3\" | \"F4\" | \"F5\",\n \"rationale\": \"one sentence explaining the verdict\"\n}\n\nIf verdict is \"correct\", failure_mode MUST be null.\nIf verdict is \"incorrect\", failure_mode MUST be one of F1-F5."}]}], "generationConfig": {"temperature": 0.0, "maxOutputTokens": 512}}}
|
|
{"request": {"contents": [{"role": "user", "parts": [{"text": "You are evaluating whether an LLM's answer is correct against ground truth.\n\n## Question\nWhen did Melanie paint a sunrise?\n\n## Ground-truth answer\n2022\n\n## Ground-truth supporting context (excerpt shown to the model)\nSession 1 (1:56 pm on 8 May, 2023):\nMelanie: You'd be a great counselor! Your empathy and understanding will really help the people you work with. By the way, take a look at this. [image: a photo of a painting of a sunset over a lake]\n\n## Model's answer\nunknown\n\n## Your task\n\nStep 1: Determine if the model's answer is correct.\n- \"correct\" means the model's answer contains all required facts from ground truth, with no additional incorrect claims.\n- Minor phrasing differences, synonyms, or alternative but equivalent formulations are acceptable.\n- Extra detail is acceptable ONLY if it is factually correct.\n\nStep 2: If incorrect, assign exactly one failure mode using this decision tree:\n\n1. Does the model explicitly refuse or say it does not know? ->F1 (ABSTAIN)\n2. Does the model answer a DIFFERENT question than was asked (coherent but off-topic)? ->F5 (OFF-TOPIC)\n3. Does the model rely on entities, names, dates, or claims that do NOT appear in the ground-truth context (fabrication)? ->F4 (HALLUCINATED)\n4. Does the model correctly state SOME required facts but miss others, without stating any incorrect facts? ->F2 (PARTIAL)\n5. Otherwise (model states facts derived from the context but gets them wrong): ->F3 (INCORRECT)\n\nStep 3: Return JSON only, no prose, in this exact schema:\n\n{\n \"verdict\": \"correct\" | \"incorrect\",\n \"failure_mode\": null | \"F1\" | \"F2\" | \"F3\" | \"F4\" | \"F5\",\n \"rationale\": \"one sentence explaining the verdict\"\n}\n\nIf verdict is \"correct\", failure_mode MUST be null.\nIf verdict is \"incorrect\", failure_mode MUST be one of F1-F5."}]}], "generationConfig": {"temperature": 0.0, "maxOutputTokens": 512}}}
|
|
{"request": {"contents": [{"role": "user", "parts": [{"text": "You are evaluating whether an LLM's answer is correct against ground truth.\n\n## Question\nWhat fields would Caroline be likely to pursue in her educaton?\n\n## Ground-truth answer\nPsychology, counseling certification\n\n## Ground-truth supporting context (excerpt shown to the model)\nSession 1 (1:56 pm on 8 May, 2023):\nCaroline: Gonna continue my edu and check out career options, which is pretty exciting!\nCaroline: I'm keen on counseling or working in mental health - I'd love to support those with similar issues.\n\n## Model's answer\nunknown\n\n## Your task\n\nStep 1: Determine if the model's answer is correct.\n- \"correct\" means the model's answer contains all required facts from ground truth, with no additional incorrect claims.\n- Minor phrasing differences, synonyms, or alternative but equivalent formulations are acceptable.\n- Extra detail is acceptable ONLY if it is factually correct.\n\nStep 2: If incorrect, assign exactly one failure mode using this decision tree:\n\n1. Does the model explicitly refuse or say it does not know? ->F1 (ABSTAIN)\n2. Does the model answer a DIFFERENT question than was asked (coherent but off-topic)? ->F5 (OFF-TOPIC)\n3. Does the model rely on entities, names, dates, or claims that do NOT appear in the ground-truth context (fabrication)? ->F4 (HALLUCINATED)\n4. Does the model correctly state SOME required facts but miss others, without stating any incorrect facts? ->F2 (PARTIAL)\n5. Otherwise (model states facts derived from the context but gets them wrong): ->F3 (INCORRECT)\n\nStep 3: Return JSON only, no prose, in this exact schema:\n\n{\n \"verdict\": \"correct\" | \"incorrect\",\n \"failure_mode\": null | \"F1\" | \"F2\" | \"F3\" | \"F4\" | \"F5\",\n \"rationale\": \"one sentence explaining the verdict\"\n}\n\nIf verdict is \"correct\", failure_mode MUST be null.\nIf verdict is \"incorrect\", failure_mode MUST be one of F1-F5."}]}], "generationConfig": {"temperature": 0.0, "maxOutputTokens": 512}}}
|
|
{"request": {"contents": [{"role": "user", "parts": [{"text": "You are evaluating whether an LLM's answer is correct against ground truth.\n\n## Question\nWhat did Caroline research?\n\n## Ground-truth answer\nAdoption agencies\n\n## Ground-truth supporting context (excerpt shown to the model)\nSession 2 (1:14 pm on 25 May, 2023):\nCaroline: Researching adoption agencies — it's been a dream to have a family and give a loving home to kids who need it.\n\n## Model's answer\nunknown\n\n## Your task\n\nStep 1: Determine if the model's answer is correct.\n- \"correct\" means the model's answer contains all required facts from ground truth, with no additional incorrect claims.\n- Minor phrasing differences, synonyms, or alternative but equivalent formulations are acceptable.\n- Extra detail is acceptable ONLY if it is factually correct.\n\nStep 2: If incorrect, assign exactly one failure mode using this decision tree:\n\n1. Does the model explicitly refuse or say it does not know? ->F1 (ABSTAIN)\n2. Does the model answer a DIFFERENT question than was asked (coherent but off-topic)? ->F5 (OFF-TOPIC)\n3. Does the model rely on entities, names, dates, or claims that do NOT appear in the ground-truth context (fabrication)? ->F4 (HALLUCINATED)\n4. Does the model correctly state SOME required facts but miss others, without stating any incorrect facts? ->F2 (PARTIAL)\n5. Otherwise (model states facts derived from the context but gets them wrong): ->F3 (INCORRECT)\n\nStep 3: Return JSON only, no prose, in this exact schema:\n\n{\n \"verdict\": \"correct\" | \"incorrect\",\n \"failure_mode\": null | \"F1\" | \"F2\" | \"F3\" | \"F4\" | \"F5\",\n \"rationale\": \"one sentence explaining the verdict\"\n}\n\nIf verdict is \"correct\", failure_mode MUST be null.\nIf verdict is \"incorrect\", failure_mode MUST be one of F1-F5."}]}], "generationConfig": {"temperature": 0.0, "maxOutputTokens": 512}}}
|
|
{"request": {"contents": [{"role": "user", "parts": [{"text": "You are evaluating whether an LLM's answer is correct against ground truth.\n\n## Question\nWhat is Caroline's identity?\n\n## Ground-truth answer\nTransgender woman\n\n## Ground-truth supporting context (excerpt shown to the model)\nSession 1 (1:56 pm on 8 May, 2023):\nCaroline: The transgender stories were so inspiring! I was so happy and thankful for all the support. [image: a photo of a dog walking past a wall with a painting of a woman]\n\n## Model's answer\nunknown\n\n## Your task\n\nStep 1: Determine if the model's answer is correct.\n- \"correct\" means the model's answer contains all required facts from ground truth, with no additional incorrect claims.\n- Minor phrasing differences, synonyms, or alternative but equivalent formulations are acceptable.\n- Extra detail is acceptable ONLY if it is factually correct.\n\nStep 2: If incorrect, assign exactly one failure mode using this decision tree:\n\n1. Does the model explicitly refuse or say it does not know? ->F1 (ABSTAIN)\n2. Does the model answer a DIFFERENT question than was asked (coherent but off-topic)? ->F5 (OFF-TOPIC)\n3. Does the model rely on entities, names, dates, or claims that do NOT appear in the ground-truth context (fabrication)? ->F4 (HALLUCINATED)\n4. Does the model correctly state SOME required facts but miss others, without stating any incorrect facts? ->F2 (PARTIAL)\n5. Otherwise (model states facts derived from the context but gets them wrong): ->F3 (INCORRECT)\n\nStep 3: Return JSON only, no prose, in this exact schema:\n\n{\n \"verdict\": \"correct\" | \"incorrect\",\n \"failure_mode\": null | \"F1\" | \"F2\" | \"F3\" | \"F4\" | \"F5\",\n \"rationale\": \"one sentence explaining the verdict\"\n}\n\nIf verdict is \"correct\", failure_mode MUST be null.\nIf verdict is \"incorrect\", failure_mode MUST be one of F1-F5."}]}], "generationConfig": {"temperature": 0.0, "maxOutputTokens": 512}}}
|