Logging & Tracing
How to get structured logs, distributed traces, and per-run debugging data out of Agent Builder workflow executions.
Structured logs
Every workflow-run, node, and tool-call log line carries run_id, workflow_id, node_id, and node_type as structured fields — set by RunLogger at the run/node level and propagated down to built-in tool handlers (api_caller, sql_query, etc.) via the same context, so a failure inside a tool call can be traced straight back to the run and node that triggered it.
| Env var | Purpose |
|---|---|
LOG_FORMAT=json | Emit JSON-formatted log lines (default: plain text) |
GLOBAL_LOG_LEVEL | Log level, e.g. INFO, DEBUG, WARNING |
Example JSON log line from a failed tool call:
{
"level": "WARNING",
"message": "_resolve_secret failed for 'stripe_key': secret not found",
"run_id": "run_abc123",
"workflow_id": "wf_xyz",
"node_id": "node_7",
"node_type": "agent"
}Outside of an active workflow run (e.g. a tool handler invoked directly in a unit test), these fields are simply omitted — logging falls back to the plain module logger.
Distributed tracing (OpenTelemetry)
Agent Builder is instrumented with OpenTelemetry. When enabled, each workflow run produces a trace with this hierarchy:
workflow.run (one span per run)
└── workflow.node (one span per node)
└── <auto-instrumented spans> (httpx / aiohttp / sqlalchemy calls made
by that node's tool calls or MCP/API
requests — nested under the node span
that triggered them)
| Env var | Purpose | Default |
|---|---|---|
ENABLE_OTEL | Master switch | false |
ENABLE_OTEL_TRACES | Enable trace export | false |
ENABLE_OTEL_METRICS | Enable metrics export | false |
ENABLE_OTEL_LOGS | Enable log export via OTLP (in addition to stdout) | false |
OTEL_EXPORTER_OTLP_ENDPOINT | Collector endpoint | http://localhost:4317 |
OTEL_SERVICE_NAME | Service name shown in your trace backend | hrida-ai-studio |
Point OTEL_EXPORTER_OTLP_ENDPOINT at any OTLP-compatible collector (Jaeger, Tempo, Datadog agent, etc.) to view traces.
A node that calls a built-in tool — api_caller hitting an external REST API, the mcp node calling an MCP server, sql_query against an external database — shows that call's span nested under the workflow.node span for the node that made it, not as a disconnected trace. This lets you see, for a slow node, exactly how much of its latency was the tool call itself vs. everything else the node did.
Decision-path attributes
if_else, while_loop, classify, and transform node spans carry a decision.* attribute set (e.g. decision.condition, decision.result, decision.raw_response, decision.matched_class) recording what was actually evaluated and why that branch/class was taken — so "why did the workflow take this path" is reconstructible from the trace itself, not just the resulting branch label.
Metrics
Emitted under the hridaai.workflow.* and hridaai.llm.* namespaces (requires ENABLE_OTEL_METRICS=true):
| Metric | Type | Description |
|---|---|---|
hridaai.workflow.runs.total | Counter | Total runs, tagged by final status |
hridaai.workflow.run.duration | Histogram | End-to-end run duration (ms) |
hridaai.workflow.node.duration | Histogram | Per-node execution duration (ms) |
hridaai.workflow.tokens.total | Counter | LLM tokens consumed |
hridaai.workflow.cost.usd | Histogram | Estimated LLM cost per run |
hridaai.workflow.runs.active | Up-down counter | Runs currently in-flight |
hridaai.llm.ttft | Histogram | Time-to-first-token per LLM call |
hridaai.llm.errors.total | Counter | LLM call errors, tagged by provider/type |
hridaai.workflow.timeout_ceiling_hits.total | Counter | Runs stopped by the wall-clock WORKFLOW_MAX_TIMEOUT_SECONDS ceiling — distinct from an ordinary node failure |
A single ceiling hit can be a genuinely long-running case. A rising rate of hridaai.workflow.timeout_ceiling_hits.total for one workflow usually means a stuck while_loop/hung LLM call, or a ceiling set too low for that workflow's real workload — worth alerting on separately from the general run.duration histogram, since it flags runs that were cut off rather than ones that simply finished slowly.
Querying a run's timeline
Beyond logs and traces, every run persists a per-event timeline (node starts, outputs, completions, errors) that you can query directly — useful for building a run-replay UI or debugging without a tracing backend configured.
GET /api/v1/agent-workflows/{workflow_id}/runs/{run_id}/timelineGET /api/v1/agent-analytics/runs/{run_id}/timelineNode latency stats (P50/P95 per node type, computed from timeline data):
GET /api/v1/agent-analytics/node-statsSee Run Infrastructure for the run-history and analytics endpoints, and Analytics for workflow-level cost/usage reporting.