Skip to main content

Logging & Tracing

How to get structured logs, distributed traces, and per-run debugging data out of Agent Builder workflow executions.


Structured logs​

Every workflow-run, node, and tool-call log line carries run_id, workflow_id, node_id, and node_type as structured fields — set by RunLogger at the run/node level and propagated down to built-in tool handlers (api_caller, sql_query, etc.) via the same context, so a failure inside a tool call can be traced straight back to the run and node that triggered it.

Env varPurpose
LOG_FORMAT=jsonEmit JSON-formatted log lines (default: plain text)
GLOBAL_LOG_LEVELLog level, e.g. INFO, DEBUG, WARNING

Example JSON log line from a failed tool call:

{
  "level": "WARNING",
  "message": "_resolve_secret failed for 'stripe_key': secret not found",
  "run_id": "run_abc123",
  "workflow_id": "wf_xyz",
  "node_id": "node_7",
  "node_type": "agent"
}
info

Outside of an active workflow run (e.g. a tool handler invoked directly in a unit test), these fields are simply omitted — logging falls back to the plain module logger.


Distributed tracing (OpenTelemetry)​

Agent Builder is instrumented with OpenTelemetry. When enabled, each workflow run produces a trace with this hierarchy:

workflow.run                     (one span per run)
└── workflow.node (one span per node)
└── <auto-instrumented spans> (httpx / aiohttp / sqlalchemy calls made
by that node's tool calls or MCP/API
requests — nested under the node span
that triggered them)
Env varPurposeDefault
ENABLE_OTELMaster switchfalse
ENABLE_OTEL_TRACESEnable trace exportfalse
ENABLE_OTEL_METRICSEnable metrics exportfalse
ENABLE_OTEL_LOGSEnable log export via OTLP (in addition to stdout)false
OTEL_EXPORTER_OTLP_ENDPOINTCollector endpointhttp://localhost:4317
OTEL_SERVICE_NAMEService name shown in your trace backendhrida-ai-studio

Point OTEL_EXPORTER_OTLP_ENDPOINT at any OTLP-compatible collector (Jaeger, Tempo, Datadog agent, etc.) to view traces.

Node-level tool calls are nested correctly

A node that calls a built-in tool — api_caller hitting an external REST API, the mcp node calling an MCP server, sql_query against an external database — shows that call's span nested under the workflow.node span for the node that made it, not as a disconnected trace. This lets you see, for a slow node, exactly how much of its latency was the tool call itself vs. everything else the node did.

Decision-path attributes​

if_else, while_loop, classify, and transform node spans carry a decision.* attribute set (e.g. decision.condition, decision.result, decision.raw_response, decision.matched_class) recording what was actually evaluated and why that branch/class was taken — so "why did the workflow take this path" is reconstructible from the trace itself, not just the resulting branch label.


Metrics​

Emitted under the hridaai.workflow.* and hridaai.llm.* namespaces (requires ENABLE_OTEL_METRICS=true):

MetricTypeDescription
hridaai.workflow.runs.totalCounterTotal runs, tagged by final status
hridaai.workflow.run.durationHistogramEnd-to-end run duration (ms)
hridaai.workflow.node.durationHistogramPer-node execution duration (ms)
hridaai.workflow.tokens.totalCounterLLM tokens consumed
hridaai.workflow.cost.usdHistogramEstimated LLM cost per run
hridaai.workflow.runs.activeUp-down counterRuns currently in-flight
hridaai.llm.ttftHistogramTime-to-first-token per LLM call
hridaai.llm.errors.totalCounterLLM call errors, tagged by provider/type
hridaai.workflow.timeout_ceiling_hits.totalCounterRuns stopped by the wall-clock WORKFLOW_MAX_TIMEOUT_SECONDS ceiling — distinct from an ordinary node failure
Watch the timeout-ceiling counter as a leading indicator

A single ceiling hit can be a genuinely long-running case. A rising rate of hridaai.workflow.timeout_ceiling_hits.total for one workflow usually means a stuck while_loop/hung LLM call, or a ceiling set too low for that workflow's real workload — worth alerting on separately from the general run.duration histogram, since it flags runs that were cut off rather than ones that simply finished slowly.


Querying a run's timeline​

Beyond logs and traces, every run persists a per-event timeline (node starts, outputs, completions, errors) that you can query directly — useful for building a run-replay UI or debugging without a tracing backend configured.

GET /api/v1/agent-workflows/{workflow_id}/runs/{run_id}/timeline
GET /api/v1/agent-analytics/runs/{run_id}/timeline

Node latency stats (P50/P95 per node type, computed from timeline data):

GET /api/v1/agent-analytics/node-stats

See Run Infrastructure for the run-history and analytics endpoints, and Analytics for workflow-level cost/usage reporting.

Hrida.ai is proprietary software of Zlabs Innovation. See the license for terms. © 2026 Zlabs Innovation.