What it does
The run went through TypeSafe's playground rather than a script, so what is here is the questions and the raw answers. Its own reading is that low confidence overlaps with the wrong answers.
Benchmarks & research
The 2026 Korean CSAT Korean-language paper turned into a JSON state and typed choice questions, with Jev's answers, probabilities and error analysis kept alongside.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
The run went through TypeSafe's playground rather than a script, so what is here is the questions and the raw answers. Its own reading is that low confidence overlaps with the wrong answers.
You can test how well an AI handles complex multiple-choice tests. Start by gathering your reading passages and five-choice questions with their answer keys. A developer can help format them into structured files that separate the background passages from the specific answer options.
Submit the questions through TypeSafe to see the predicted answers and certainty scores. Notice that questions relying on visual diagrams or layout clues lose critical information in plain text. You should also check low-certainty scores carefully, as they frequently signal wrong answers.