What it does
Will Rice tests this approach on six BFCL subsets. He used test examples during development, so the reported scores are not clean held-out results or official leaderboard placements.
Benchmarks & research
Let Jev pick a tool and the values it needs, then have code assemble the call.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Will Rice tests this approach on six BFCL subsets. He used test examples during development, so the reported scores are not clean held-out results or official leaderboard placements.
Ask a developer to give your assistant a list of tools, including a choice to use none. After Jev picks a tool, offer the possible numbers, dates or other values it needs. Code puts the choices together. This project tests proposed calls; it does not carry out the actions.
Trying new requests sends their text, any earlier-message context and tool descriptions to TypeSafe's outside AI service, using paid credits. Share only permitted material. The author used some test examples during development; he did not test conversations over several turns or calls to several tools at once.