<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Agentic AI | Anh Duong Vo</title><link>https://anhduongvo.github.io/tags/agentic-ai/</link><atom:link href="https://anhduongvo.github.io/tags/agentic-ai/index.xml" rel="self" type="application/rss+xml"/><description>Agentic AI</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>https://anhduongvo.github.io/media/icon_hu_982c5d63a71b2961.png</url><title>Agentic AI</title><link>https://anhduongvo.github.io/tags/agentic-ai/</link></image><item><title>Agentic AI tooling and evaluation</title><link>https://anhduongvo.github.io/projects/agentic-tooling/</link><pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate><guid>https://anhduongvo.github.io/projects/agentic-tooling/</guid><description>&lt;p&gt;Open tools that make agentic AI easier to build, trust and evaluate. These are framework and vendor neutral, which is the point: the hard parts of agentic AI (discovery, evaluation, grounding) are the same whatever model you run.&lt;/p&gt;
&lt;h2 id="tools"&gt;Tools&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;mcp-bionemo.&lt;/strong&gt; A Model Context Protocol server that exposes NVIDIA BioNeMo NIMs (RFdiffusion, ProteinMPNN, Boltz-2) as typed tools, so any MCP client can call them. BioNeMo ships as agent skills and raw endpoints but not as MCP, so this closes a real gap. Runs on a simulator by default, one flag switches to the live NIMs.
&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;clinical-agent-eval.&lt;/strong&gt; A reusable harness that scores clinical agent outputs on grounding, number accuracy, hallucinated-citation rate and calibration, and renders a leaderboard.
&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;
(framework-agnostic).&lt;/strong&gt; The same evaluation ideas generalized across agent frameworks (LangGraph, LlamaIndex, OpenAI tool-calling), so one harness scores traces from any of them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;
.&lt;/strong&gt; Retrieval-augmented generation over public clinical guidelines and drug labels, with every answer verified against its cited source, on a neutral open stack.&lt;/p&gt;
&lt;h2 id="live-demo"&gt;Live demo&lt;/h2&gt;
&lt;p&gt;An interactive app (offline, no key needed) that runs the verification logic of the clinical agents:
&lt;/p&gt;</description></item><item><title>Clinical Agentic AI (open examples)</title><link>https://anhduongvo.github.io/projects/clinical-agentic-ai/</link><pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate><guid>https://anhduongvo.github.io/projects/clinical-agentic-ai/</guid><description>&lt;p&gt;Four open, end-to-end examples of agentic AI for healthcare and the life sciences, built on the open NVIDIA stack (NeMo Agent Toolkit, NIM, Nemotron, NeMo Guardrails, BioNeMo). They run on synthetic and public data only, so the implementations can be shared and reproduced without any protected data.&lt;/p&gt;
&lt;p&gt;One design principle runs through all of them: in healthcare, a fluent wrong answer is the real risk. So the model cites its evidence, code checks every number, the model judges only meaning, and a human signs off.&lt;/p&gt;
&lt;h2 id="the-four-examples"&gt;The four examples&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;
, ambient clinical documentation.&lt;/strong&gt; Turns a consultation transcript into a structured clinical note in which every sentence cites the evidence it rests on, code verifies doses and numbers, and a clinician accepts or rejects each item. Speech recognition uses Riva/Parakeet, and the project ships in batch, live, and served modes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;
, clinical-trial screening.&lt;/strong&gt; Checks a patient&amp;rsquo;s FHIR record against a study&amp;rsquo;s eligibility criteria and marks each criterion met, not met, or unknown, with the facts it rests on and a calibrated confidence. Rules decide age, sex, and lab thresholds in code; the model handles the rest; a coordinator confirms before anyone is contacted.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;
, regulatory writing.&lt;/strong&gt; Checks a draft clinical study report against the ICH E3 structure, drafts missing sections from the source tables with a row citation per sentence, and verifies every number in code against the cited row. A medical writer accepts or rejects each finding.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;
, protein binder design.&lt;/strong&gt; For a target protein, writes a cited literature brief, then chains the BioNeMo models (RFdiffusion, ProteinMPNN, Boltz-2) to design, fold, and score candidate binders into a ranked, honest report. It runs on a deterministic simulator by default; one flag switches each step to the real BioNeMo NIMs.&lt;/p&gt;
&lt;h2 id="stack"&gt;Stack&lt;/h2&gt;
&lt;p&gt;NeMo Agent Toolkit, NIM, Nemotron, NeMo Guardrails, Riva/Parakeet ASR, guided JSON (constrained decoding), BioNeMo NIMs, FHIR, Python.&lt;/p&gt;
&lt;p&gt;All four are open source on GitHub:
.&lt;/p&gt;
&lt;h2 id="try-it-and-build-on-it"&gt;Try it and build on it&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Live demo.&lt;/strong&gt; An interactive app (offline, no key needed) that runs the verification logic of all four examples:
.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Related tooling.&lt;/strong&gt;
, a Model Context Protocol server that exposes the BioNeMo NIMs (RFdiffusion, ProteinMPNN, Boltz-2) as tools, and
, a reusable harness that scores grounding, number accuracy, hallucinated citations and calibration across the four projects.&lt;/p&gt;</description></item></channel></rss>