<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>RAG | Anh Duong Vo</title><link>https://anhduongvo.github.io/tags/rag/</link><atom:link href="https://anhduongvo.github.io/tags/rag/index.xml" rel="self" type="application/rss+xml"/><description>RAG</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>https://anhduongvo.github.io/media/icon_hu_982c5d63a71b2961.png</url><title>RAG</title><link>https://anhduongvo.github.io/tags/rag/</link></image><item><title>Agentic AI tooling and evaluation</title><link>https://anhduongvo.github.io/projects/agentic-tooling/</link><pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate><guid>https://anhduongvo.github.io/projects/agentic-tooling/</guid><description>&lt;p&gt;Three tools around the clinical agents, written so that they work with any model or framework. Each has a short terminal recording; everything in the recordings runs offline without an API key.&lt;/p&gt;
&lt;h2 id="1-mcp-bionemo-bionemo-models-as-mcp-tools"&gt;1. mcp-bionemo: BioNeMo models as MCP tools&lt;/h2&gt;
&lt;p&gt;BioNeMo&amp;rsquo;s biology models are available as NeMo Agent Toolkit agent skills and as HTTP endpoints, but not as a Model Context Protocol server, so MCP clients such as Claude Desktop, Cursor or an IDE agent cannot discover or call them directly. This server wraps the RFdiffusion, ProteinMPNN and Boltz-2 endpoints in typed MCP tools (&lt;code&gt;design_backbone&lt;/code&gt;, &lt;code&gt;design_sequences&lt;/code&gt;, &lt;code&gt;fold_complex&lt;/code&gt;, and a one-call &lt;code&gt;design_binder&lt;/code&gt;). It runs on a deterministic simulator by default and switches to the real NIMs with one environment variable, including the asynchronous job polling the hosted biology NIMs use.&lt;/p&gt;
&lt;p&gt;The recording (21 s) runs the tool tests against the simulator and shows the client configuration that registers the server.&lt;/p&gt;
&lt;video controls poster="/projects/agentic-tooling/mcp-bionemo.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/agentic-tooling/mcp-bionemo.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="2-agenteval-agent-evaluation-at-two-levels"&gt;2. agenteval: agent evaluation at two levels&lt;/h2&gt;
&lt;p&gt;Level 1 scores how an agent behaves, from traces. Adapters normalise LangGraph, LlamaIndex and OpenAI tool-calling traces into one schema, and the metrics cover tool success, tool selection, grounding of citations in retrieved context, task success, mean steps, and repeated tool calls. Level 2 scores what the agent says: grounding rate, hallucinated-citation rate, number accuracy with rounding tolerance, and calibration (ECE and Brier), with a per-task leaderboard. The clinical agents emit the level 2 schema, so one harness scores all of them.&lt;/p&gt;
&lt;p&gt;The recording (22 s) runs both levels on the bundled samples; the claims leaderboard picks up a planted wrong number and a planted citation that does not exist.&lt;/p&gt;
&lt;video controls poster="/projects/agentic-tooling/agenteval.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/agentic-tooling/agenteval.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="3-rag-guidelines-retrieval-with-citation-verification"&gt;3. rag-guidelines: retrieval with citation verification&lt;/h2&gt;
&lt;p&gt;Retrieves the relevant passages from a corpus of clinical guidelines and drug labels, answers only from those passages with a citation on every sentence, and then checks each sentence against the chunk it cites: enough overlap in content words, and every number present in the source. Sentences that fail are flagged. It uses any OpenAI-compatible endpoint (vLLM, Ollama, TGI, hosted APIs) and has an offline extractive mode that needs no key. The bundled corpus is synthetic.&lt;/p&gt;
&lt;p&gt;The recording (21 s) answers two questions and shows the retrieved sources with their scores and the per-sentence check.&lt;/p&gt;
&lt;video controls poster="/projects/agentic-tooling/rag-guidelines.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/agentic-tooling/rag-guidelines.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="code"&gt;Code&lt;/h2&gt;
&lt;p&gt;
·
·
·
&lt;/p&gt;</description></item><item><title>Clinical Agentic AI (open examples)</title><link>https://anhduongvo.github.io/projects/clinical-agentic-ai/</link><pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate><guid>https://anhduongvo.github.io/projects/clinical-agentic-ai/</guid><description>&lt;p&gt;Four open, end-to-end examples of agentic AI for healthcare and the life sciences, built with the NVIDIA NeMo stack (NeMo Agent Toolkit, NIM, Nemotron, NeMo Guardrails, BioNeMo). They run on synthetic and public data only, so the code can be shared and reproduced without any protected data.&lt;/p&gt;
&lt;p&gt;All four follow the same pattern: the model drafts, code checks the numbers and citations, and a person reviews the result before it is used. Each example below has a short video recorded from the interactive demo. The checks in the videos run live; the generated clinical text is a bundled sample.&lt;/p&gt;
&lt;h2 id="1-consult-to-note-ambient-clinical-documentation"&gt;1. consult-to-note: ambient clinical documentation&lt;/h2&gt;
&lt;p&gt;Turns a consultation transcript into a structured SOAP note. Each sentence of the note cites the transcript lines it came from, numbers such as doses and vitals are compared against those lines in code, and a clinician accepts or rejects each sentence before the note is exported as a FHIR document. Speech recognition uses Riva/Parakeet with boosting for drug names, and the pipeline runs in batch mode, live during the consultation, and as a served endpoint.&lt;/p&gt;
&lt;p&gt;The video (36 s) shows two consultations. For each, the note is generated with its numbers verified, then an error is planted (250 mg instead of 25 mg; 25 units instead of 2.5 units) and the check flags the sentence against the transcript line.&lt;/p&gt;
&lt;video controls poster="/projects/clinical-agentic-ai/consult-to-note.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/clinical-agentic-ai/consult-to-note.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="2-trial-matcher-clinical-trial-screening"&gt;2. trial-matcher: clinical-trial screening&lt;/h2&gt;
&lt;p&gt;Checks a patient&amp;rsquo;s FHIR record against the eligibility criteria of a ClinicalTrials.gov study. Each criterion gets a status (met, not met, or unknown) together with the FHIR resources it rests on and a confidence value. Age, sex and lab thresholds are evaluated in code; criteria that need reading are passed to the model; a lab value that is missing or older than a year is reported as unknown. A study coordinator reviews the result before the patient is contacted.&lt;/p&gt;
&lt;p&gt;The video (33 s) shows two patients. For the first, the eGFR criterion is unknown until a value is added, and lowering HbA1c flips that criterion to not met. For the second, two criteria start as not met and are resolved by editing the record.&lt;/p&gt;
&lt;video controls poster="/projects/clinical-agentic-ai/trial-matcher.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/clinical-agentic-ai/trial-matcher.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="3-csr-assistant-clinical-study-report-review"&gt;3. csr-assistant: clinical study report review&lt;/h2&gt;
&lt;p&gt;Checks a draft clinical study report against the ICH E3 section structure, drafts missing sections from the source tables with a row citation on every sentence, and reviews the text: every number is compared in code against the row it cites, and qualitative claims such as &amp;ldquo;well tolerated&amp;rdquo; are checked against the data by the model. A medical writer accepts or rejects each finding.&lt;/p&gt;
&lt;p&gt;The video (29 s) shows two sections. In the efficacy section a placebo-arm change of -1.4 is flagged against the table&amp;rsquo;s -0.35 and verified once corrected; in the safety section a discontinuation count of 3 is flagged against the table&amp;rsquo;s 2.&lt;/p&gt;
&lt;video controls poster="/projects/clinical-agentic-ai/csr-assistant.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/clinical-agentic-ai/csr-assistant.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="4-ai-scientist-protein-binder-design"&gt;4. ai-scientist: protein binder design&lt;/h2&gt;
&lt;p&gt;For a target protein, the agent writes a literature brief that cites only retrieved abstracts, then chains three BioNeMo models: RFdiffusion designs backbones for the epitope, ProteinMPNN designs sequences for each backbone, and Boltz-2 co-folds each binder with the target and scores the complex. Candidates are ranked by a composite score into a report that states its method and limits. The project runs on a deterministic simulator by default; one flag switches each step to the real BioNeMo NIMs.&lt;/p&gt;
&lt;p&gt;The video (25 s) changes the target, the random seed and the number of candidates and shows the ranking update. The scores come from the simulator and are placeholders.&lt;/p&gt;
&lt;video controls poster="/projects/clinical-agentic-ai/ai-scientist.jpg" &gt;
&lt;source src="https://anhduongvo.github.io/projects/clinical-agentic-ai/ai-scientist.mp4" type="video/mp4"&gt;
&lt;/video&gt;
&lt;h2 id="stack"&gt;Stack&lt;/h2&gt;
&lt;p&gt;NeMo Agent Toolkit, NIM, Nemotron, NeMo Guardrails, Riva/Parakeet ASR, guided JSON (constrained decoding), BioNeMo NIMs, FHIR, Python.&lt;/p&gt;
&lt;p&gt;Code:
,
,
,
.&lt;/p&gt;
&lt;h2 id="related"&gt;Related&lt;/h2&gt;
&lt;p&gt;The interactive app the videos were recorded from, which runs offline without a key:
. An MCP server for the BioNeMo models:
. An evaluation harness that scores the four projects&amp;rsquo; claims and agent behavior:
.&lt;/p&gt;</description></item></channel></rss>