Juris ranked #2 globally on IBM VAKRA for tool selection
Our agent Juris ranked #2 globally on the IBM VAKRA benchmark specifically for tool selection. Not a bad way to end the week.
The benchmark
VAKRA evaluates how well agents select and use tools in complex enterprise workflows. The tool space is roughly 8,000+ APIs across about 62 domains.
Doing less, choosing better
A lot of work went into making the agent do less and choose better. A 36B model turned out to be more than enough; larger models tended to introduce more failures, timeouts and unnecessary token usage.
For VAKRA capability 2, Juris did not run a generic multi-step ReAct loop. We used a capability-specific routing stack: shortlist the tool space, force a single best tool choice, normalise arguments for brittle endpoints, and return deterministic answers directly from tool outputs instead of letting the model paraphrase and drift.
Learn more: IBM’s VAKRA announcement, Juris at ClerkTree Research and the VAKRA leaderboard on Hugging Face.
