<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation | Anh Duong Vo</title><link>https://anhduongvo.github.io/tags/evaluation/</link><atom:link href="https://anhduongvo.github.io/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><description>Evaluation</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>https://anhduongvo.github.io/media/icon_hu_982c5d63a71b2961.png</url><title>Evaluation</title><link>https://anhduongvo.github.io/tags/evaluation/</link></image><item><title>Agentic AI tooling and evaluation</title><link>https://anhduongvo.github.io/projects/agentic-tooling/</link><pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate><guid>https://anhduongvo.github.io/projects/agentic-tooling/</guid><description>&lt;p&gt;Three tools around the clinical agents, written so that they work with any model or framework. Each has a short terminal recording; everything in the recordings runs offline without an API key.&lt;/p&gt;
&lt;h2 id="1-mcp-bionemo-bionemo-models-as-mcp-tools"&gt;1. mcp-bionemo: BioNeMo models as MCP tools&lt;/h2&gt;
&lt;p&gt;BioNeMo&amp;rsquo;s biology models are available as NeMo Agent Toolkit agent skills and as HTTP endpoints, but not as a Model Context Protocol server, so MCP clients such as Claude Desktop, Cursor or an IDE agent cannot discover or call them directly. This server wraps the RFdiffusion, ProteinMPNN and Boltz-2 endpoints in typed MCP tools (&lt;code&gt;design_backbone&lt;/code&gt;, &lt;code&gt;design_sequences&lt;/code&gt;, &lt;code&gt;fold_complex&lt;/code&gt;, and a one-call &lt;code&gt;design_binder&lt;/code&gt;). It runs on a deterministic simulator by default and switches to the real NIMs with one environment variable, including the asynchronous job polling the hosted biology NIMs use.&lt;/p&gt;
&lt;p&gt;The recording (21 s) runs the tool tests against the simulator and shows the client configuration that registers the server.&lt;/p&gt;
&lt;video controls poster="/projects/agentic-tooling/mcp-bionemo.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/agentic-tooling/mcp-bionemo.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="2-agenteval-agent-evaluation-at-two-levels"&gt;2. agenteval: agent evaluation at two levels&lt;/h2&gt;
&lt;p&gt;Level 1 scores how an agent behaves, from traces. Adapters normalise LangGraph, LlamaIndex and OpenAI tool-calling traces into one schema, and the metrics cover tool success, tool selection, grounding of citations in retrieved context, task success, mean steps, and repeated tool calls. Level 2 scores what the agent says: grounding rate, hallucinated-citation rate, number accuracy with rounding tolerance, and calibration (ECE and Brier), with a per-task leaderboard. The clinical agents emit the level 2 schema, so one harness scores all of them.&lt;/p&gt;
&lt;p&gt;The recording (22 s) runs both levels on the bundled samples; the claims leaderboard picks up a planted wrong number and a planted citation that does not exist.&lt;/p&gt;
&lt;video controls poster="/projects/agentic-tooling/agenteval.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/agentic-tooling/agenteval.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="3-rag-guidelines-retrieval-with-citation-verification"&gt;3. rag-guidelines: retrieval with citation verification&lt;/h2&gt;
&lt;p&gt;Retrieves the relevant passages from a corpus of clinical guidelines and drug labels, answers only from those passages with a citation on every sentence, and then checks each sentence against the chunk it cites: enough overlap in content words, and every number present in the source. Sentences that fail are flagged. It uses any OpenAI-compatible endpoint (vLLM, Ollama, TGI, hosted APIs) and has an offline extractive mode that needs no key. The bundled corpus is synthetic.&lt;/p&gt;
&lt;p&gt;The recording (21 s) answers two questions and shows the retrieved sources with their scores and the per-sentence check.&lt;/p&gt;
&lt;video controls poster="/projects/agentic-tooling/rag-guidelines.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/agentic-tooling/rag-guidelines.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="code"&gt;Code&lt;/h2&gt;
&lt;p&gt;
·
·
·
&lt;/p&gt;</description></item></channel></rss>