Tools

AI-generated text

Using NeMo Relay to Trace and Evaluate Hermes Agent Behavior

NVIDIA’s NeMo Relay integration with Hermes Agent provides structured traces (ATOF, ATIF, and OpenTelemetry) to inspect model and tool calls, errors, retries, timing, and token usage.

Using NeMo Relay to Trace and Evaluate Hermes Agent Behavior

An agent can complete a task and still take inefficient or redundant steps: repeated searches, re-reading truncated files, or extra model calls. Those inefficiencies raise latency and token consumption and increase the chance of failures. To improve an agent’s behavior developers need more than a binary success check—they need structured evidence of how the task was executed.

This tutorial demonstrates two Hermes Agent examples run with NVIDIA NeMo Relay and uses the resulting traces to inspect model and tool calls, errors, retries, durations, and token usage. It shows how to combine that trace evidence with task verification, and it includes a Hermes ToolPerf benchmark case that illustrates harness evaluation across repeated runs.

What the tutorial covers

The text and accompanying video explain how to:

  • Set up an isolated Hermes Agent runtime with native NeMo Relay integration.
  • Run a simple terminal-tool task and inspect the event stream and trajectory.
  • Run a file-and-web research task and explore its OpenTelemetry trace in Arize Phoenix.
  • Combine task verification with trace evidence to evaluate a harness change.

Prerequisites

Before beginning you need:

  • macOS or Linux
  • Git and curl
  • Docker Desktop or Docker Engine running
  • An NVIDIA Build API key for NVIDIA Nemotron 3.5 Lightning (use the model page’s Generate API Key action)

How NeMo Relay integrates with the Hermes Agent harness

NeMo Relay provides a common way to observe and control model and tool execution. The Hermes Agent harness includes NeMo Relay natively: sessions, turns, model calls, and tool calls are represented in Relay’s scope hierarchy, and NeMo Relay records lifecycle events (start and end) with timing and parent-child relationships preserved.

Trace outputs and when to use them

The tutorial works with three complementary representations of agent execution:

  • Agent Trajectory Observability Format (ATOF): a JSONL log of scope starts, scope ends, and point-in-time marks with IDs and timestamps. Use ATOF to debug or audit individual events, timing, and parent-child relationships.
  • Agent Trajectory Interchange Format (ATIF): a step-by-step JSON record assembled from lifecycle events that describes the agent’s interactions and tool calls. Use ATIF to review or analyze the agent’s path step by step.
  • OpenTelemetry (with the OpenInference exporter): records the run as parent-child spans and labels agent, LLM, and tool spans with attributes. Use OTEL traces in OTEL-compliant tools like Arize Phoenix to inspect model and tool calls, duration, token usage, and errors.

An ATIF tool request shows what the model asked to run, but it does not confirm the tool’s outcome; verify by locating the matching tool start and end events (and any errors) in the ATOF. Shared uuid values pair related events; parent_uuid connects a tool call to its parent scope.

Review traces before sharing: depending on configuration they may include prompts, model responses, tool arguments and results, file paths, and other application data. NeMo Relay provides an evidence layer that enterprises and auditors can use to investigate agent behavior, evaluate policies, and extend security controls.

Experiment #1: run a simple tool-use task with Hermes Agent

The first example is intentionally small and deterministic. Hermes uses its terminal tool to run a bundled Python script inside an isolated Docker container. The script prints:

VALUE=42

That fixed output provides an exact success check. A passing run also confirms Hermes reached the model, invoked the terminal tool in the sandbox, and produced Relay trace files. The container has no network, repository checkout, or NVIDIA API key access; Hermes cannot fall back to running host terminal commands.

To reproduce, run these commands and add your NVIDIA_API_KEY to keys.env after copying the template:

  • git clone https://github.com/NVIDIA/nemoclaw-community
  • cd nemoclaw-community/examples/tools/hermes-relay-tracing
  • ./scripts/setup_tutorial_runtime.sh
  • cp keys.env.example keys.env (add NVIDIA_API_KEY to keys.env)
  • docker version (verify Docker is running)
  • ./scripts/build_tutorial_image.sh
  • ./scripts/run_tutorial.sh

The setup script creates a self-contained environment under .tutorial-runtime/ with dependencies such as Python 3.11, Hermes 0.21.1, and NeMo Relay 0.8.3. It does not alter your existing Python or Hermes installation. Docker is used for the terminal-tool sandbox and the local Phoenix service.

When the task finishes, the runner verifies the response and the generated traces. A successful run prints a verification line and ATOF/ATIF summaries. Example output from one verified run (values vary by run):

ATOF summary

  • events: 74
  • completed llm scopes: 2
  • llm scopes with usage: 2
  • prompt tokens: 7239
  • completion tokens: 96
  • total tokens: 7335
  • tool calls: 1
  • tool errors: 0
  • correlated events: 74

ATIF summary

  • agent: Hermes Agent
  • model: nvidia/nemotron-3.5-lightning-30b-a3b
  • steps: 3
  • llm calls: 2
  • requested tool calls: 1

Task verified: VALUE=42

The Artifacts path identifies the run directory that contains the complete ATOF event stream and ATIF trajectory. The companion repository describes how to inspect or summarize either file further.

Experiment #2: run a multi-tool research task and explore traces in Phoenix

The second example reuses the same Hermes and NeMo Relay environment for a task that requires multiple tools. Hermes receives a travel record with clues about an unnamed machine-learning conference. The agent must read the record, find a conference matching subject, dates, and location, confirm the answer on the official site, save a verified report, and return the conference name.

Use NeMo Relay’s OpenInference exporter to send OpenTelemetry spans to Arize Phoenix over OTLP. Phoenix displays an interactive trace where you can inspect model and tool calls, timing, token usage, errors, and inputs/outputs.

You can send the same OTEL trace to other OTLP-compatible backends (for example LangSmith) by changing endpoint and authentication. The NeMo Relay observability guide documents exporters and configuration options.

This example reuses the NVIDIA Nemotron model and the NVIDIA API key from experiment #1. Hermes uses its keyless web search and the runner starts Phoenix in a pinned local container. Run the conference search example:

./scripts/run_conference_research_with_phoenix.sh

During the run, NeMo Relay saves ATOF and ATIF locally. Before reporting success the runner checks that Hermes:

  • identified COLT 2026 as the conference name;
  • saved a report containing the expected conference details and an official source;
  • successfully completed read_file, web_search, web_extract, and write_file calls;
  • produced a nonempty ATIF trajectory;
  • sent model and tool spans with positive token usage to Phoenix.

After the checks pass, the terminal prints a link to the Phoenix project and the local run directory. Open the project in Phoenix to follow the run from the initial file read through web search, source verification, report write, and final response.

Note: to compare model behavior, follow the companion repository’s model-profile instructions and keep query, tools, execution limits, and verifier constant. Because the task uses live web search, these runs are best for exploring behavior rather than definitive model ranking; controlled comparisons require fixed search responses and repeated runs.

Using traces to evaluate a harness change

The two examples show verification and single-run inspection. Evaluating a harness change requires the same checks under controlled, repeated conditions. Recommended steps:

  1. Choose a fixed task with an exact, automated success check.
  2. Define a baseline and a single focused change to prompt, tool, configuration, or harness.
  3. Keep everything else constant (model snapshot, provider, task input, execution budget, timeout).
  4. Run the same number of repetitions for baseline and candidate with NeMo Relay enabled.
  5. Compare verified outcomes first, then use traces to examine model calls, tool calls, retries, errors, elapsed time, token usage, and cost.
  6. Repeat across the models and workloads the change is expected to affect before generalizing.

Call a candidate an improvement only when it produces a repeatable increase in completion or preserves completion while improving reliability, latency, or cost. A single faster run or fewer calls may explain a result but does not establish an optimization by itself. An unchanged result can be informative if it shows dependence on model, environment, or a particular sample.

Benchmark case study: evaluating Hermes tool-layer changes (Hermes ToolPerf)

Nous Research developed the Hermes ToolPerf benchmark by analyzing production sessions, auditing tool schemas, and mining logs for failure classes. They used NeMo Relay ATOF traces as ground truth for turn accounting, mapped nine failure patterns to deterministic benchmark cases, and used those cases to evaluate a batch of Hermes tool-layer fixes.

The August 6 rerun compared a pinned baseline and fixed revisions across nine tasks. Each task ran three times per model per arm, totaling 108 runs with the same prompts, tools, execution limits, and success checks. A task verifier measured completion and NeMo Relay ATOF traces captured model calls, tool calls, errors, retries, tool-result data, and timing.

Summary results (two models, baseline vs fixes):

  • Claude Sonnet 4.5 — Baseline: 24/27 (89%) success; mean LLM calls 2.9; mean tool calls 2.2; mean tool result data 17 KB; mean duration 16 s. Fixes: 23/27 (85%); 2.8 LLM calls; 2.1 tool calls; 17 KB; 22 s.
  • Qwen3 Coder 30B — Baseline: 19/27 (70%) success; mean LLM calls 3.8; mean tool calls 2.8; mean tool result data 16 KB; mean duration 27 s. Fixes: 22/27 (81%); mean LLM calls 4.9; mean tool calls 3.9; mean tool result data 33 KB; mean duration 42 s.

Interpretation: Sonnet showed no meaningful change between baseline and fixes. Qwen3 Coder improved completion (three more successful tasks) but required more LLM and tool calls, larger tool-result payloads, and longer durations—indicating the fixes enabled recoveries on tasks the baseline abandoned, at the cost of a chattier, slower agent in some cases. Task-level audit traced these trade-offs to specific failure patterns (for example, recovery recipes that increased turns when the baseline had given up, or extra exploratory searches that regressed performance).

The NeMo Relay traces provided the detailed evidence—every model call, tool call, error, retry, result payload, and timing was recorded for all 108 runs—and the raw trace files are checked into the results directory for reproducibility.

Getting started

NeMo Relay offers a consistent way to capture evidence for evaluating agent harness changes. ATOF preserves ordered lifecycle events; ATIF presents the same work as a readable trajectory; OTEL exporters enable interactive inspection in tools like Arize Phoenix. Pair these traces with deterministic verifiers to compare harness changes without conflating fewer calls with better outcomes.

The companion repository contains the runnable tasks, verifiers, NeMo Relay configuration, and Phoenix setup used in the tutorial. See the NeMo Relay and Hermes Agent documentation, ATOF/ATIF docs, and the NeMo Relay observability configuration guide for installation and configuration details.