Observability¶
Guide: how the built-in distributed tracing works, and how to export it to OpenTelemetry.
At a glance
- What — every traced request produces a causal span tree across the three layers; the system stamps W3C-compatible trace/span IDs automatically and can ship the telemetry to any OpenTelemetry backend.
- Built-in — distributed tracing is part of the platform; you turn it on per endpoint/flow and the spans propagate without code.
- For developers and operators wiring Mercury to Dynatrace, Splunk, Jaeger, Tempo, or an OpenTelemetry Collector.
- Logs too — an opt-in application log context stamps the same correlation/trace ids (plus your own key-values) into structured log lines, so logs and spans join up in your backend.
In an event-driven, composable system a single request fans out across decoupled functions, flows, and graph nodes that never call each other directly. Observability is therefore not optional — it is the only way to see the causal path a request actually took. Mercury answers this with a built-in distributed-tracing engine whose output is OpenTelemetry-compliant, so the spans drop straight into the dashboard you already run.
The built-in tracing design¶
Turning tracing on¶
Tracing is opt-in per entry point:
- HTTP endpoints — add
tracing: trueto therest.yamlentry (see REST Automation). - Event Script flows — a flow started through
FlowExecutorcarries a trace; the engine traces every task. - Programmatically — construct a trace-aware
PostOffice(fromRoute, traceId, tracePath)and the events it sends are traced end-to-end.
Two controls tune what is recorded:
@ZeroTracingon a function suppresses tracing for that function (used by system services so they never trace themselves).skip.rpc.tracing(inapplication.properties) lists route names excluded from trace recording; the default isasync.http.request.
What a trace records¶
When a traced function finishes, the system sends a performance-metrics dataset to the built-in
distributed.tracing service. The dataset is a plain map:
trace={ id=<32-hex trace id>, span_id=<16-hex>, parent_span_id=<16-hex>,
service=<route name>, path=<request path>, from=<caller route>,
origin=<instance id>, start=<ISO-8601>, exec_time=<ms>, round_trip=<ms>,
success=<bool>, status=<int>, exception=<text> }
annotations={ <your key>=<value>, ... }
span_id and parent_span_id are the OpenTelemetry-compatible IDs (see W3C trace context); you can attach
business context to the span with PostOffice.annotateTrace(key, value). To attach context to application logs
instead, see Application log context.
Spans across the three layers¶
The span tree mirrors the three paradigm layers, and tracing is virtual-thread-safe — the
parent span is threaded through a per-task anchor, not a ThreadLocal, so concurrently-dispatched siblings still
share the correct parent.
| Layer | What becomes a span | Lineage |
|---|---|---|
| Layer 1 — Platform Core | each function execution | the caller's span becomes the child's parent_span_id |
| Layer 2 — Event Script | each task, plus one synthetic task.executor flow-summary span (annotated with the flow id) |
tasks chain exactly like Layer 1; a sub-flow chains to the parent task that dispatched it |
| Layer 3 — Knowledge Graph | each node dispatch through graph.executor |
the graph traversal threads the parent span through node execution |
Because Layer 2 and Layer 3 ride on Layer 1's engine, the task/node spans look identical to Layer-1 spans — the only addition at Layer 2 is the synthetic flow summary that brackets the whole flow's timing.
W3C Trace Context (OpenTelemetry compliance)¶
Mercury's trace and span IDs follow the W3C Trace Context format that
OpenTelemetry uses: a 32-hex trace ID and 16-hex span ID. Across an HTTP boundary the system propagates the
standard traceparent header:
- outbound — the HTTP client injects
traceparentcarrying the current trace + span, alongsideX-Trace-Id; - inbound — the HTTP layer extracts
traceparent, continues the upstream trace, and adopts the caller's span as the parent.traceparenttakes precedence overX-Trace-Idwhen both are present.
The X-Trace-Id header carries the trace ID for callers not yet on W3C Trace Context. The framework does not
echo the trace ID back to the HTTP client. (The correlation-id — a separate concern — is documented in
Reserved Names & Headers.)
Let the framework manage trace headers — don't set them yourself. Inside a traced flow or function the platform injects
X-Trace-Idandtraceparenton every outbound HTTP call from the current trace context, overwriting any value you set on the request — so the trace stays correct and connected end-to-end, and the upstream trace is propagated automatically. The one exception is a call that is not being traced (an endpoint withtracing: false, or a call made outside a trace): there a trace header you set passes through untouched — the intended escape hatch for handing a trace context to a third-party system, or for unit-testing an external endpoint with full control over its request headers.
Header impedance matching (trace-id and correlation-id)¶
Not every caller names its headers the way this framework does. An enterprise gateway may have standardized on
its own trace header long before W3C Trace Context; a legacy Kafka producer may stamp X-Correlation-ID; and
two systems bridged by one application rarely agree with each other. Rather than forcing every party to rename,
the platform matches the impedance at the edge: the header names are configuration, while everything
downstream keeps working with the same two ids —
- the trace id (with its W3C
traceparentcontext) drives the distributed tracing on this page; - the business correlation-id is a separate concern — captured at the edge, preserved as the flow's
model.cid, and exposed to every function viaPostOffice.getMyCorrelationId()(see Reserved Names & Headers).
The configurable names¶
Global defaults in application.properties:
| Key | Default | Where it applies |
|---|---|---|
http.trace.id.header |
X-Trace-Id |
REST automation inbound (when no traceparent is present) and the async HTTP client outbound |
http.correlation.id.header |
X-Correlation-Id |
HTTP edge capture inbound; async HTTP client outbound |
kafka.trace.id.header |
(unset) | Kafka Flow Adapter inbound fallback; simple.kafka.notification outbound, stamped alongside traceparent |
kafka.correlation.id.header |
cid |
Kafka Flow Adapter inbound; simple.kafka.notification outbound |
Per-entry overrides, for a single application that faces callers with different conventions:
- a rest.yaml endpoint entry accepts
trace.id.header/correlation.id.header; - a kafka-flow-adapter.yaml consumer binding accepts the same two keys (Minimalist Kafka);
- twin-kafka adds
secondary.kafka.trace.id.header/secondary.kafka.correlation.id.headerglobals when the second Kafka cluster follows its own convention.
Precedence: per-entry override > application.properties global > built-in default. A well-formed W3C
traceparent always takes precedence for the trace id, whatever the header naming — so adopting these
overrides never breaks OpenTelemetry-compliant callers. The full key reference lives in the
Configuration Reference.
Legacy conflation is supported. Pointing the trace-id name at the correlation-id header (
http.trace.id.header=X-Correlation-Id) is a valid backward-compatibility setup for an estate whose gateway only passes that one header. The edge keeps the two ids consistent: a supplied shared header feeds both ids, and when the shared header is absent the edge resolves one id for both — from the inboundtraceparentwhen present, otherwise a single generated id — so the outgoingtraceparentand the shared header always carry the same trace id. Prefer migrating the gateway to passtraceparent(andX-Trace-Id), then retiring the conflation.
# rest.yaml - one endpoint serves a legacy caller that sends its own header names
- service: "legacy.orders"
methods: ['POST']
url: "/api/legacy/orders"
timeout: 15s
tracing: true
trace.id.header: "X-Legacy-Trace"
correlation.id.header: "X-Legacy-Cid"
Bridging two conventions¶
When one application connects two systems that each own a correlation-id convention (for example a dual-cluster Kafka bridge), keep each system's header name strictly on its own side:
- the inbound adapter binding declares that system's
correlation.id.header, so the id lands inmodel.cid; - the flow maps it back out under the next system's name —
'model.cid -> header.X-Their-Header'— on the outbound data mapping; - neither system ever sees the other's header name: the id value is what crosses the bridge, and the trace
context (
traceparent) rides alongside automatically, keeping one continuous distributed trace end to end.
The twin-kafka guide covers this pattern across two Kafka clusters, and the
twin-kafka-demo worked
example runs it end to end: an on-prem system's X-Correlation-Id and a cloud system's
X-Cloud-Correlation-Id carry the same id through one bridged transaction, with each header name confined to
its own cluster.
Exporting telemetry¶
The forwarder extension point¶
By default the trace dataset is logged. To ship it elsewhere, register a function at the reserved route
distributed.trace.forwarder; the system detects it and forwards every dataset to it. A companion hook,
transaction.journal.recorder, receives request/response payloads when journaling is enabled (see
Reserved Names & Headers and Build, Test & Deploy).
The OpenTelemetry forwarder (ready-made)¶
You do not have to write the forwarder for OpenTelemetry. The opentelemetry-forwarder extension ships a
distributed.trace.forwarder that maps each dataset to an OpenTelemetry span — preserving the exact W3C
trace/span/parent-span IDs — and exports it over OTLP/HTTP to a collector (and on to Dynatrace, Splunk, Jaeger,
Tempo, …). Add the dependency and it auto-registers; no code required:
Note:
x.y.zdenotes the current Mercury version shown in the rootpom.xml.
<dependency>
<groupId>org.platformlambda</groupId>
<artifactId>opentelemetry-forwarder</artifactId>
<version>x.y.z</version>
</dependency>
Configure it in application.properties (values support ${ENV_VAR:default} substitution):
otel.exporter.otlp.endpoint=${OTEL_EXPORTER_OTLP_ENDPOINT:http://localhost:4318/v1/traces}
otel.service.name=${OTEL_SERVICE_NAME:my-app}
# Backend credentials come from the environment — no secret hard-coded:
otel.exporter.otlp.headers=${OTEL_EXPORTER_OTLP_HEADERS}
Point the endpoint at an OpenTelemetry Collector, or directly at a SaaS backend with its API token in the headers:
# Dynatrace
export OTEL_EXPORTER_OTLP_ENDPOINT="https://{env-id}.live.dynatrace.com/api/v2/otlp/v1/traces"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Api-Token dt0c01.XXXX"
# Splunk Observability Cloud
export OTEL_EXPORTER_OTLP_ENDPOINT="https://ingest.{realm}.signalfx.com/v2/trace/otlp"
export OTEL_EXPORTER_OTLP_HEADERS="X-SF-Token=YOUR_ACCESS_TOKEN"
Each span carries the route name (route), path, from, origin, status, timing, and your annotation.*
values as span attributes, with service.name on the resource. See the full key reference in the
Configuration Reference.
A custom forwarder¶
To target a system without an OTLP path, implement your own function at distributed.trace.forwarder and consume
the dataset described in What a trace records — for example, write it to a metrics database or a
proprietary APM API.
Application log context¶
Spans tell you the causal path; application logs tell you what happened inside each step. The log-context
feature closes the gap between them: it injects a context block — correlation id, trace/span ids, service name,
and any business key-values you add — into every structured log line a traced function emits. With the same
traceId/spanId on both the span and the log line, you can pivot from a Dynatrace/Splunk trace straight to the
exact log entries that belong to it.
It deliberately avoids the ThreadLocal / Log4j MDC pattern (heavy for a virtual-thread runtime). The context rides the same per-request mechanism as the trace itself, keyed to the worker thread and torn down when the function returns.
Turning it on¶
The feature is opt-in and activates only when an optional app-log-context.yaml is on the classpath
(src/main/resources/). It applies to the two structured JSON appenders — select one via the log4j2
configuration (log4j2-json.xml for pretty output, log4j2-compact.xml for single-line); the plain Console
appender is unaffected.
# src/main/resources/app-log-context.yaml
context:
cid: $cid
traceId: $traceId
tracePath: $tracePath
spanId: $spanId
parentSpanId: $parentSpanId
service: $service
timestamp: $utc
environment: '${ENV_NAME:dev}'
hello: world
The left side is the output key (your choice). The right side is one of three forms:
| Form | Example | Resolved |
|---|---|---|
Reserved $token |
service: $service |
live, per log line, from the request's trace context |
${ENV:default} substitution |
environment: '${ENV_NAME:dev}' |
once at startup, from the environment |
| Hardcoded literal | hello: world |
emitted verbatim on every line |
The reserved tokens are $cid, $traceId, $tracePath, $spanId, $parentSpanId, $service (the current
function's route), and $utc (the log line's UTC timestamp). A token (or env value) that resolves to nothing is
omitted from the block rather than printed as null — so a root span simply has no parentSpanId key.
Adding your own key-values¶
Inside a function, add business context with PostOffice.updateContext(key, value):
var po = new PostOffice(headers, instance);
po.updateContext("user", "demo"); // appears in the context block of every subsequent log line
log.info("processing request");
The reserved keys (cid, traceId, tracePath, spanId, parentSpanId, service, utc) are protected —
passing one to updateContext throws IllegalArgumentException. On a non-traced request, or when the feature is
off, the call is a silent no-op.
updateContextvsannotateTrace— two distinct sinks.annotateTrace(...)attaches business data to the distributed-trace dataset that flows to your APM backend (What a trace records);updateContext(...)attaches it to the application log stream only. Neither leaks into the other.
What it looks like¶
A log line from a traced function then carries the resolved context (from the worked example above, with
po.updateContext("user", "demo") added in the function):
{
"level": "INFO",
"context": {
"cid": "20260630c6ee70d866cb4fae9ab3c44d926ce21a",
"traceId": "fbb60df209084531b2b00f6b36a3e651",
"tracePath": "GET /api/profile/100",
"spanId": "bf8d4b2b6a923d67",
"parentSpanId": "98f8e26ae7d9a422",
"service": "v1.hello.exception",
"environment": "dev",
"hello": "world",
"user": "demo",
"timestamp": "2026-06-30T21:17:03Z"
},
"time": "2026-06-30 14:17:03.575",
"source": "com.accenture.demo.tasks.HelloExceptionHandler.handleEvent(HelloExceptionHandler.java:51)",
"thread": 297,
"message": "User defined exception handler - status=404 error=Profile 100 not found"
}
The traceId and spanId here match the v1.hello.exception span the tracer emitted for the same request, so the
log line and the span join up in your backend. (Key order within context is not significant — log viewers reorder
keys on display.)
Scope and boundaries¶
- The
contextblock appears only when a request is traced and atraceIdis present. Framework boot logs and logs emitted from aMono/Fluxcompletion that runs after the worker returns (on a different thread) carry no context — the same boundary distributed tracing has. - Feature off (no
app-log-context.yaml) costs one boolean check per log line and nothing else.
See also¶
- Build, Test & Deploy — enabling tracing and the forwarder hook.
- Configuration Reference — the
trace.*andotel.*keys. - Reserved Names & Headers —
distributed.trace.forwarder,traceparent. - Architecture Overview — the three layers the span tree mirrors.