Skip to main content

Observability System

Overview​

Chronon provides a unified observability system with three integrated capabilities:

The same unit APIs support both implicit single-clock simulations and explicit hardware clock domains. Explicit clocks retain domain/local-edge attribution and use exact rational physical time for ordering. The native ClockTraceRecorder remains available for clockEvent and CDC protocol events, including alongside unified observation. See explicit clock semantics before interpreting counter intervals or comparing scheduler and model timelines.

FeaturePurposeAPIHot Path
CountersStatistics collection++counter_ / counter_ += n~1-2ns
Timeline eventsStructured event captureevent<"name">(CAT, ...)fixed record + typed args
Pipeline tracesTyped one-cycle pipeline slicesmodel-level observe::pipeline<"STAGE">(...)fixed record + typed args
LogsDebug outputdebug<"fmt">(...)~2ns disabled

Explicit clock domains​

event, debug/info, EventCounter, named spans and pipeStage retain their existing signatures. Configure observation before initialize(), then register counters and start the backend after initialization. SimulationApp manages this lifecycle. Direct C++ callers can use:

observe::ObservationYAMLConfig observation;
observation.enabled = true;
observation.unified_logging.trace_channel.enabled = true;
observation.unified_logging.categories.push_back({"my_category", true, {}});
observation.counters.reference_clock = "cpu";
observation.counters.periodic_dump_cycles = 100;
sim.configureObservation(observation);
sim.initialize();
auto& manager = observe::ObservationManager::instance();
manager.reregisterAllCounters();
manager.startBackend();
sim.runUntilTime(SimTime::nanoseconds(10'000));
sim.finalize();
manager.dumpFinalCounterSnapshot(0); // Clock mode gets its cutoff from sim.
sim.writeTimelineTrace();
manager.stopBackend();
API/outputExplicit-clock semantics
Structured events and spansDomain ID, local edge and exact ClockDomain::edge(n) time; typed arguments, strings and flow IDs are preserved.
pipeStageOccupancy [edge(n), edge(n+1)) in the owning hardware domain.
Text logsdomain, edge, and exact rational time prefix. Category time filters use that source's local edge index.
Periodic countersA shared physical sampling grid defined by reference_clock and periodic_dump_cycles; default means the configured default tick frequency, regardless of unit domains.
Scheduler diagnosticsHost monotonic nanoseconds; separate scheduler_timeline.pftrace. Cluster slices include domain, local edge and physical time in their detail. Migration logging never reads the coordinator's non-atomic calendar state.
Native clock recordingExisting configureClockTrace interface, identities, protocol flows and output files remain unchanged.

The unified backend converts exact time to floor(seconds * 1e9) only for Perfetto display. Structured records also export domain_id, local_cycle, time_num and time_den, so distinct subnanosecond edges remain inspectable. Text timestamps and counter CSV retain exact rational seconds. Source identity and equal-time model-event ordering follow logical unit names, independent of worker placement and migration. Lifecycle records are ordered before/after model records at their boundary; same-phase lifecycle callbacks retain publication order. A worker may interleave unrelated hardware domains; its largest observed local cycle is never used as a global watermark.

Each counter owner captures and resets its own counters before executing the first edge at or after a periodic cutoff. A slow/idle owner also services cutoffs reached by retired physical progress without executing extra model ticks. Rows therefore cover [previous cutoff, cutoff) even when the nominal cutoff falls between a domain's edges. Counter ownership and the next sampling index follow stable scheduler clusters across migration. Sequential execution keeps distinct owners for units from different domains.

Clock CSV begins with time_num,time_den,sample. periodic rows share the reference grid. A successful runUntilTime(t) final snapshot is final_before at its exclusive cutoff t, including a no-work call before the first phased edge. Batch/domain limits and early termination use final_after at the last committed edge: this includes the work at that exact time. These phases keep before-edge and after-edge samples distinct when timestamps coincide. Repeated final dumps without new work are ignored; a final dump resets the residual interval, and subsequent runs retain the original periodic sampling phase. Legal no-work calls with an earlier or equal cutoff preserve the last observation boundary and its before/after phase, including for later finalization. Stopping and resuming preserves the scheduler's existing bounded settlement contract. Observations add no model ticks or CDC acceptance decisions.

initialize() observations occur at time zero; finalize() observations use the last run's actual cutoff. Such structured events carry a lifecycle annotation, and logs label the callback explicitly. Their local_cycle is the unit's next edge index, not a claim that this edge executed. Finalize before the final counter dump to include counter changes made by that callback. If a caller has already dumped the same run cutoff, finalization permits one additional residual row at that cutoff; earlier counter values are not emitted again. At the same exact time, a pre-finalization snapshot precedes the finalize callback records, and the post-finalization residual follows them. This lifecycle ordering is independent of the hardware cutoff: CSV still reports final_before or final_after according to the last run.

The backend reuses producer queues, source/format/track registries, counter plans, the reorder arena and isolated HostServices I/O lane. A rolling credit window bounds admitted observation timestamps. Before advancing the safe frontier it captures all queue heads after acquiring scheduler progress, then drains that complete prefix. Queue pressure obeys the existing channel policy; the backend never flushes an unsafe prefix merely to free space. If one open window exceeds service_buffer_bytes, it reports an error rather than silently misordering data. Increase that budget or reduce recording rate/lookahead for bursty models. Continuous drain and producer assistance prevent a quiet stream or several actors sharing a worker from blocking ingress draining.

Unsupported combinations fail explicitly: starting the unified backend before clock initialization, unknown counter reference domains, speculative observation epochs attached to explicit clocks, pipeline occupancy in lifecycle callbacks, and submitting host-wall streams into the physical-time writer. Use a separate scheduler output file. Observation records emitted behind already published physical progress also fail explicitly. Unit observation calls and counter updates remain owner-thread operations; external threads must not mutate a running unit's observation state.

sender_test_multiclock_observation exercises a phased 3 GHz/400 MHz FIFO model with existing APIs, serial/epoch-free/migrating state equivalence, segmented runs, termination/resume, native tracing alongside unified tracing, lifecycle callbacks, bounded-buffer saturation, backend failure and recovery. Real Perfetto import, flow, exact-time and interval-counter checks are available with:

python3 scripts/validate_multiclock_observation.py \
--trace-processor /path/to/trace_processor \
--binary build/test/sender/sender_test_multiclock_observation

Recording every edge incurs queueing, sorting, formatting, serialization and I/O costs. Compare disabled, counter-only and full-recording workloads separately; these APIs do not make per-edge tracing free.

To reproduce a short overhead check with state equivalence outside the timer:

build/test/sender/sender_test_multiclock_observation --overhead /tmp/clock-observe-cost

The fixture runs three FIFO pairs for 3 µs (30,600 unit edges), emitting one structured event, pipeline slice and text log per edge in full mode. Counter mode samples every 128 fast-domain edges; native recording is off. Seven interleaved trials per mode on an Intel i9-14900K, GCC Release build, measured the following wall times including final snapshots and backend drain (initialization excluded):

WorkersDisabled median (range), msCounters median (range), msFull median (range), ms
10.796 (0.762–3.929)4.316 (4.146–8.188)39.068 (37.478–40.654)
32.248 (2.206–2.305)7.341 (7.124–7.685)42.126 (40.757–43.371)

These are small, very cheap model ticks; admission, interval sampling and I/O costs dominate. The shared host had unrelated CPU load and no fixed affinity, so these measurements demonstrate recording cost rather than a throughput guarantee or regression claim. Re-run against the intended model and machine.

Design Principles​

  • Cheap disabled path: Existing enable checks bypass producer and backend work
  • Minimal overhead when enabled: Pre-registered format strings
  • Lock-free hot path: No mutex contention
  • Lookahead-compatible: Thread-local counters, buffered events

Quick Start​

Define Categories and Counters​

#include "chronon/Chronon.hpp"
using namespace chronon;

// Categories - bit positions auto-assigned at program startup
// Category<> objects auto-register when constructed as global/static variables
inline const auto CACHE_HIT = Category<"cache_hit", "Cache hit events">{};
inline const auto CACHE_MISS = Category<"cache_miss", "Cache miss events">{};

Use in Units​

Units created through TickSimulation automatically use their owning unit's local cycle for event timestamps. A getObserveCycle() override is only needed for standalone observers or a deliberate custom time source. This binding does not add work to the simulation's tick dispatch.

class FetchUnit : public TickableUnit, public ObservableUnit {
EventCounter cache_hits_{this, "cache_hits", "Cache hit count"};

public:
void tick() override {
// Increment counter
++cache_hits_;

// Update the aggregate and emit a structured timeline event
cache_hits_.mark<"cache_hit">(CACHE_HIT, arg<"pc">(pc));

// Logging
debug<"Fetch cycle {} pc=0x{:x}">(localCycle(), pc);
info<"Started fetching">();
}
};

Initialize and Run​

For most users, SimulationApp handles observation setup automatically from YAML:

int main(int argc, char* argv[]) {
return chronon::SimulationApp("My Simulator")
.setDefaultConfig("config.yaml")
.run(argc, argv);
}

For manual setup without SimulationApp:

int main() {
TickSimulation sim;
auto* fetch = sim.createUnit<FetchUnit>();

auto& obs = ObservationManager::instance();
obs.initialize(yaml_config);

auto* ctx = obs.createContextForUnit(
"fetch", [&]() { return sim.currentCycle(); });
fetch->setObservationContext(ctx);

// Enable trace categories
ctx->filter().enableCategory(CACHE_HIT);
ctx->filter().setMinLogLevel(LogLevel::Debug);

obs.startBackend();

sim.initialize();
sim.run(1000);

obs.stopBackend();
obs.shutdown();
}

Log Levels​

enum class LogLevel : uint8_t {
Debug = 0, // Verbose debugging
Info = 1, // General information
Warn = 2, // Warnings
Error = 3 // Errors
};

ctx.filter().setMinLogLevel(LogLevel::Info); // Debug filtered

Category Filtering​

O(1) bitmask-based filtering:

ctx.filter().enableCategory(CACHE_HIT);
ctx.filter().disableCategory(CACHE_MISS);

if (ctx.shouldTrace(CACHE_HIT)) { /* ... */ }

Lookahead Support​

For speculative execution:

ctx.setLookaheadMode(true);

// Speculative work
counter_ += 10;
event<"speculative_event">(CAT, arg<"value">(value));

if (mispredicted) {
ctx.rollbackEpoch(); // Counters restored, events discarded
} else {
ctx.commitEpoch(); // Epoch advances, events flushed
}

ctx.setLookaheadMode(false);

Output Files​

out/
└── 20260124_143052/
├── counters.csv # Counter snapshots
├── timeline.pftrace # Perfetto protobuf timeline (unit events and counters)
├── scheduler_timeline.pftrace # Optional scheduler execution timeline
└── events.log # Text logs (debug/info/warn/error)
  • events.log holds text output for the debug/info/warn/error log channels.
  • timeline.pftrace contains unit events and counter tracks (see below); created when observation.timeline.enabled is true (the default).
  • scheduler_timeline.pftrace contains scheduler execution slices when scheduler timeline collection is enabled.
  • counters.csv is created when counters are enabled with csv_output: true.

Perfetto Timeline​

Timeline events and counter snapshots are written to timeline.pftrace in Perfetto protobuf format. Scheduler execution data uses the separate scheduler_timeline.pftrace file. Open either file directly in ui.perfetto.dev, or query it offline with Perfetto's trace_processor.

The timeline contains:

  • Timeline events and lanes — typed instants and occupancy spans emitted through ObservableUnit::event/instant/spanBegin/spanEnd, TimelineLane, or TimelineSpan. Hierarchical unit paths become nested track groups.
  • Counter tracks — one Perfetto counter track per counter, sampled at counter dump cycles and nested under each unit's collapsible counters subgroup (unit.counters.counter_name).
  • Scheduler execution timeline — the separate scheduler file contains wall-clock unit/wait/epoch slices under a "Chronon Scheduler" process group, with one lane per worker stream plus a scheduler lane. Slices carry cycle and detail debug annotations. See Scheduler Timeline Trace.

The file is produced by src/observe/PerfettoTraceWriter, a thin wrapper over the Perfetto SDK's protozero message writers (no tracing session or category registration — packets are written straight to the file).

Pipeline Trace Events​

For cycle-by-cycle pipeline visualization, use typed pipeline events. A model should expose a small category-binding wrapper, conventionally named observe::pipeline, on top of Chronon's typed pipeline event primitive. The wrapper keeps the call site semantic while the queued record stays structured and numeric:

// Instruction uid is the item id; rendered as a one-cycle slice named "1234".
observe::pipeline<"DEC">(*this,
observe::pipeSlot(lane),
inst.uid,
observe::arg<"pc">(inst.pc),
observe::arg<"op">(inst.opcode));

// Program counter is the item id; stored as uint64_t, rendered as hex in Perfetto.
observe::pipeline<"BP0">(*this,
observe::pc(fetch_pc),
observe::arg<"seq">(fetch_seq));

// Runtime pipe/lane selection is explicit and typed.
observe::pipeline<"IF0">(*this,
observe::pipeSlot(if_pipe),
observe::pc(fetch_pc),
observe::arg<"src">(source_id));

The recommended wrapper shape is:

namespace my_model::observe {
using namespace chronon::observe;

inline const auto PIPE = Category<"pipe", "Pipeline visualization events">{};

struct PipeSlot { uint16_t value = 0; };
struct PipelineItem { uint64_t value = 0; bool hex_name = false; };

constexpr PipeSlot pipeSlot(uint64_t value) noexcept {
return {static_cast<uint16_t>(value)};
}
constexpr PipelineItem uid(uint64_t value) noexcept { return {value, false}; }
constexpr PipelineItem item(uint64_t value) noexcept { return {value, false}; }
constexpr PipelineItem pc(uint64_t value) noexcept { return {value, true}; }

// Dispatch pipeSlot.value to pipeStage<N, Stage>(...) or pipeStageHex<N, Stage>(...).
template <FixedString Stage, typename Unit, typename Id, typename... Args>
void pipeline(Unit& unit, PipeSlot pipe, Id id, Args&&... args);

template <FixedString Stage, typename Unit, typename Id, typename... Args>
void pipeline(Unit& unit, Id id, Args&&... args); // defaults to pipe 0
}

This API is intended for pipeline state; use structured member events for rare point events and debug/info/warn/error for text logs. It has these properties:

  • Structured hot path — stage and pipe are compile-time/template state or a small runtime integer; the record carries a numeric item id plus typed annotations. No backend text parsing is required for new pipeline events.
  • One-cycle slices — pipeline events render as [cycle, cycle + 1) slices, matching hardware stage occupancy better than point instants.
  • Stable coloring and flow — the numeric item id is also the Perfetto flow id and color key, so the same instruction/transaction keeps the same color across units, pipes, and stages.
  • Typed annotations — details such as pc, op, seq, addr, hit, or latency are queryable debug annotations instead of substrings in the event name.
  • Display policy is explicit at the item type — use pc(value) when the numeric id should be displayed as hexadecimal. The stored id is still a number; only the Perfetto event name formatting changes.

EventCounter snapshots are kept under each unit's counters subgroup, so dense pipeline stage tracks and low-frequency counters remain separately collapsible in the Perfetto sidebar. Pipeline lanes sort before normal timeline lanes through explicit Perfetto sibling ordering.

Wire format​

The writer uses the size-oriented parts of the Perfetto data model; all of this is transparent to ui.perfetto.dev and trace_processor:

  • Custom cycle clock — simulation events are stamped on a sequence-scoped incremental clock (clock id 64, declared via ClockSnapshot and paired 1:1 with the boot clock), so packets carry small varint cycle deltas instead of absolute timestamps. Out-of-order events (e.g. a reorder-buffer force flush) fall back to an absolute timestamp on that packet rather than corrupting the incremental state. The scheduler execution timeline stays on the default wall-clock sequence.
  • Interning — event names, categories, and debug-annotation names are emitted once per sequence as InternedData and referenced by id afterwards. Incremental state is checkpointed periodically (and whenever an intern table hits its cap) so traces stay seekable and crash-truncation-safe.
  • Compression — flushed packet batches are wrapped in zlib-deflated TracePacket.compressed_packets (on by default; timeline.compress: false disables it). Microarchitecture traces are highly repetitive, so this is where most of the size win comes from: on the bundled CPU pipeline example the timeline shrinks from ~115 MB to ~17 MB (~7×) with no measurable change in simulation wall time in that measurement. Encoding now runs on the scheduler I/O lane; performance depends on the workload and CPU allocation.

Output failures​

Trace write, flush, and close failures (for example, a full filesystem) report the output path. The backend retains its first failure and releases producers waiting on full queues. ObservationManager::stopBackend() and shutdown() rethrow the failure after completing output jobs; shutdown() also releases the manager's resources before throwing. SimulationApp returns a nonzero status when trace output fails.

Code that owns an ObservationBackend directly should call stop() followed by rethrowIfFailed(). stop() and destructors remain nonthrowing, and worker failures are also logged. Restarting the backend clears the previous failure. A standalone PerfettoTraceWriter throws from the write-triggering event, flush(), or close(); close a failed writer before reopening it. Its bytesWritten() counts only fully successful flushes, excluding a partially failed batch. It does not imply fsync durability.

Timeline Lanes and Events​

Attaching a counter-only observation context does not register lane metadata when the trace channel or timeline events are disabled. Enabling recording later registers the pending tracks at that quiescent configuration point, so event emission needs no lazy-registration work. Existing track IDs stay stable across recording toggles.

Track metadata is shared across repeated sessions with the same topology (source ID, track name, lane count, layout, and declaration attachment order within each context). Same-name lane members remain separate tracks, including those on different units sharing one context. Metadata and IDs remain valid for queued events and cached template API IDs for the process lifetime; storage scales with distinct topology declarations, not the number of simulation runs. Generating new track names or topologies indefinitely can still grow this registry.

For microarchitecture state that occupies something over many cycles — MSHR entries, ROB/LSQ slots, DRAM requests in flight, busy functional units — declare timeline lanes as unit members (no macros or registration calls):

class LSU : public Unit, public ObservableUnit {
// One sub-lane per slot; renders as a track group "mshr" with
// children mshr[0..7] under this unit's track.
TimelineLane mshr_{this, "mshr", /*lanes=*/8};
TimelineLane ld_port_{this, "ld_port", 2};
inline static const auto MISS = Category<"dcache_miss", "D$ miss lifetime">{};

void tick() override {
// Span addressed by (lane, slot): begin and end are separate calls
// and may land in different tick() invocations — no RAII scopes.
mshr_.begin(slot, MISS, "miss"_ev, flow(instr.uid),
arg<"addr">(paddr), arg<"set">(set));
...
mshr_.end(slot); // possibly many cycles later
}
};

For common model instrumentation, Chronon also provides a convenience layer that uses the same typed vocabulary as pipeline events ("name"_ev, arg<"key">(value), flow(uid)) without requiring every call site to declare raw lanes:

class Decode : public Unit, public ObservableUnit {
TimelineSpan stall_{this, "stall"};
EventCounter flushes_{this, "flushes", "Pipeline flushes", "events"};

void tick() override {
// Shared "events" track under this unit.
flushes_.mark<"flush">(FLUSH, arg<"removed">(removed_count));

// Boolean state helper: opens once, closes when the condition clears.
stall_.update(!fetch_queue_.empty() && !out_uop_queue.canSend(),
STALL, "out_uop_blocked"_ev,
arg<"fq_size">(fetch_queue_.size()),
arg<"out_rem">(out_uop_queue.remainingThisCycle()));
}
};

Use this layer for:

  • single-point events such as flushes, credit mismatches, replay markers, or rare protocol transitions (event<"flush">, instant<"track">);
  • state spans such as stalls, blocked ports, busy functional units, or waiting conditions (TimelineSpan::update, member spanBegin / spanEnd);
  • aggregate totals should use EventCounter; do not turn every counter increment into a timeline event unless the timestamp itself matters.

The vocabulary is deliberately SQL-shaped:

  • "miss"_ev — event names are low-cardinality compile-time literals (interned once in the trace), so SELECT dur FROM slice WHERE name='miss' style analysis works in trace_processor.
  • arg<"addr">(value) — per-event details go into typed debug annotations (uint/int/double/bool/pointer), not formatted into the name.
  • flow(uid) — pass the instruction/transaction uid the model already carries; Perfetto links the uid's slices across lanes and stages into one flow (click an instruction → see its whole journey), and offline analysis can join stages through the flow id to compute per-stage latency distributions. The bundled CPU pipeline example is instrumented this way: every stage stamps flow(instr_id) (fetch/dispatch/commit instants, EX occupancy spans per ALU, L2 miss spans), so examples/cpu_pipeline.yaml produces a timeline where each instruction's fetch→dispatch→ex→commit journey is one connected flow.

Semantics under the existing machinery:

  • Producers write fixed-size records to their SPSC queue without allocation; all Perfetto encoding happens on the scheduler I/O lane. With observation disabled, calls are a null-check.
  • Category and temporal filters apply to begin and instant. end() skips temporal filters so a span begun inside an observation window still closes outside it; an end whose begin was suppressed is dropped by the backend's open-span table.
  • A begin on an occupied slot implicitly closes the previous span (hardware slot reuse); spans still open at shutdown are closed at the last seen cycle.
  • Lookahead rollback discards speculative lane events; commit publishes them.

Offline analysis​

scripts/trace_sql/ ships canned trace_processor queries shaped for this data model: per-stage latency through flow edges, span-duration histograms per event name, lane occupancy, stall attribution (cycles beyond each event kind's best case), and counter statistics — all keyed by the hierarchical track paths, so they work unchanged on any model using the timeline API. See scripts/trace_sql/README.md for usage.

Structured Event Arguments​

Use a low-cardinality compile-time event name and typed arguments:

event<"request">(CAT, arg<"id">(id));

ObservationManager​

ObservationManager is the central coordinator for the observation system:

auto& obs = ObservationManager::instance();

// Initialize from YAML config
obs.initialize(yaml_config);

// Create context for each unit
auto* ctx = obs.createContextForUnit(
"fetch",
[&]() { return sim.currentCycle(); }
);
unit->setObservationContext(ctx);

// Initialize model counters and attach the backend to the simulation scheduler
sim.initialize();
obs.reregisterAllCounters();
// Start observation services and open output files
obs.startBackend();

// ... run simulation ...

// Stop backend and shutdown
obs.stopBackend();
obs.shutdown();

Key Responsibilities​

  • YAML configuration initialization: Parses config and sets up queues/backend
  • Context creation: Creates ObservationContext per unit with filtering rules
  • Backend lifecycle: Starts services, drains pending records and completes output jobs
  • Counter registration: Central registry for sparse counter pull model

ObservationQueue​

The queue is the transport layer between simulation threads and the backend:

// Constructor: capacity in bytes (rounded up to power-of-2)
ObservationQueue queue(256 * 1024); // 256 KB default

// Larger queues reduce dropped events at the cost of memory
ObservationQueue queue(1024 * 1024); // 1 MB for high-volume sims

Size Guidelines​

Simulation TypeRecommended SizeRationale
Low-volume (< 1M events/sec)256 KB (default)Minimal overhead
Medium-volume (1-10M events/sec)512 KB - 1 MBBalance memory/drops
High-volume (> 10M events/sec)2-4 MBPrevent drops during bursts

The queue uses 2x the requested capacity internally for lock-free mirroring.

Per-Thread Queues​

Trace and log events use per-thread SPSC queues to eliminate lock contention:

// Each worker thread gets its own lock-free queue
// Backend drains all queues in round-robin fashion
// Queue capacity is configured via YAML: simulation.observation.queue_capacity

Benefits:

  • No mutex contention on hot path
  • Better cache locality (thread-local writes)
  • Scales linearly with thread count

Trade-offs:

  • More memory (one queue per thread)
  • Out-of-order events (backend sees events from different threads)

The pool supports up to 64 producer slots. On thread exit, the final partial batch is published and the backend is notified. Once the consumer has drained and acknowledged that queue, a later thread can reuse its slot. Repeated thread creation therefore does not consume an ever-growing number of slots.

Queues with unread records still occupy slots after their producer exits. If all 64 slots have live producers or unread data, context acquisition fails and normal trace/log emission records a drop until a slot becomes reusable. Queue storage is cached up to the allocation high-water mark (at most 64 queues), with stable addresses and IDs for backend readers; handoff never resets consumer cursors or discards records. activeThreadCount() reports currently attached producers; allocatedContextCount() reports cached queues, including retired producers. Emission attempted by another TLS destructor after the observation producer has retired is rejected instead of creating an attachment that cannot be retired.

Scheduler service​

Both observation backends use simulation worker time for bounded ingress draining when owned by TickSimulation. For a YAML-driven single-clock simulation, the corresponding settings are:

simulation:
observation:
enabled: true
service_buffer_bytes: 16777216

For an independently owned ObservationBackend, call backend.attachScheduler(sim.hostServices()) before backend.start(). For a manually managed ObservationManager, initialize the simulation before startBackend() to share its scheduler. SimulationApp handles this order. For implicit single-clock simulations, an already-running backend retains its existing scheduler. Explicit-clock simulations require initialization before backend startup so source clocks and snapshot ownership are immutable. Native-clock simulations attach automatically through sim.configureClockTrace(config). Both sorted and immediate output use this pipeline. The old scheduler_service C++ mode switch is removed; remove the corresponding YAML key (false is rejected with a migration error, while true is accepted for existing configurations).

HostServices owns one shared I/O lane for its outputs. Each backend registers one reusable HostIOJob, with at most one outstanding batch. The executor visits pending jobs round-robin and resumes ingress only after that job's slot becomes reusable. Opening files, sorting, encoding, compression, writing, final flush and closing all run there. Standalone scheduler timeline export also uses this lane. A blocked file write keeps other I/O jobs pending and applies the configured bounded producer backpressure; it does not execute in a model worker. This lane cannot preempt an operating-system file write.

A standalone backend creates a small HostServices driver that polls the same ingress on its I/O lane. There is no separate backend consumer implementation or private backend thread pool. Its idle readiness scan sleeps up to 50 microseconds. ObservationBackend::Config::poll_interval is retained for source compatibility but is ignored; it does not change either scheduler polling or standalone waiting. New configuration members are appended to preserve existing positional aggregate initializers. Class layouts have changed, so rebuild downstream C++ binaries; binary ABI compatibility is not provided.

The scheduler visits ready host services between model sweeps and during dependency waits. Services have no simulated clock, dependency edges or lookahead frontier. A registration serializes consumers with a nonblocking claim; producer assistance uses the same registration when a synchronous tick fills its queue. Records remain owned copies, including after worker migration. Dynamic scheduling cost samples exclude time spent executing host-service polls, including producer assistance inside native-clock bridge begin/commit intervals.

Each ordinary-backend poll copies at most 256 records / 256 KiB into a preallocated handoff buffer. A native-clock poll copies at most 256 fixed-size records and preserves a snapshot of all producer heads across partial polls. Its frontier is acknowledged only after all those heads reach the I/O lane. Queue publication, progress and I/O completion signal readiness; no dedicated drain thread is created. Sorting, arena allocation, encoding, compression and file output run on the scheduler I/O lane. Serial native-clock recording signals readiness after every completed clock batch, including sparse batches and small lossy queues. This ingress notification does not advance the Perfetto watermark or close the current timestamp bucket. While I/O owns the only handoff batch, the registration rejects new claims before taking its consumer lock or reading the timing clock. Publications still record readiness, and I/O completion restores eligibility without losing notifications.

In implicit-clock mode, the ordinary reorder arena admits at most service_buffer_bytes live/retained bytes (64 KiB–1 GiB), then flushes the sorted retained records before admitting more. Its allocation may round up by less than 2×; record descriptors are separately bounded by reorder_max_events + 1. Together with the fixed handoff buffer and existing bounded producer queues, a slow sink cannot grow an unbounded raw-record backlog. Native mode accounts its handoff and head snapshot storage within the existing ingress/staging budget. Sink dictionaries, interned strings, compression buffers and scheduler timeline capture retain their own existing allocation policies; this is not a bound on whole-process memory.

As with the existing count-based forced flush, the additional ordinary-backend byte limit can flush before the reorder watermark. Ordinary-backend ordering remains best effort under forced flushing. Both explicit-clock unified observation and native clock recording instead use exact frontier ordering, without pressure-triggered flushing. Existing drop, bounded_wait and spin_wait policies are unchanged. A lossless producer can still spend host time waiting for a slow sink, while assisting bounded ingress work. No file I/O executes inside the model callback.

Stop producers and finish scheduler calls before stopping/reconfiguring a backend. Final snapshots and scheduler timelines precede detach, final drain, in-flight I/O completion and close. Detached registrations safely ignore cached assistance calls. Register custom host services only between runs and detach their registration before destroying the service object. I/O jobs retain the executor until their final completion, so an observer can close after the simulation has been destroyed. Detach ingress and wait for I/O before destroying its callback owner. Registering a job does not add model dependencies or advance simulated time.

Service statistics report poll count, records, total host elapsed nanoseconds and maximum poll duration. These are wall times (including preemption), not CPU time or hard latency deadlines. Dynamic model-cost samples exclude time spent executing service polls. Native clock-stats.json contains the same service statistics. Performance depends on CPU availability and the observation mix; compare identical worker counts, CPU affinity, outputs and loss policies.

ReorderBuffer​

The backend can reorder events by cycle for deterministic output:

ObservationBackend::Config config;
config.enable_reordering = true;
config.reorder_watermark_cycles = 1000; // Emit events older than current_cycle - 1000
config.reorder_max_events = 100000; // Max buffered events before forced flush

How It Works​

Producer threads → Per-thread queues → Backend → ReorderBuffer → Output files
↓
Sort by (cycle, source_id)
Emit when cycle < watermark

Configuration​

ParameterDefaultPurpose
enable_reorderingtrueEnable cycle-based sorting
reorder_watermark_cycles1000Cycles to buffer before emitting
reorder_max_events100000Max events before forced flush

Trade-offs:

  • Larger watermark → more deterministic, more memory
  • Smaller watermark → less memory, less out-of-order tolerance

ObservationBackend::Config​

Full configuration structure:

struct Config {
std::string output_dir = "out";
std::chrono::microseconds poll_interval{100}; // Legacy compatibility field; ignored
bool enable_counter_csv = true;
CounterCsvFormat counter_csv_format = CounterCsvFormat::Pivoted;

std::string debug_file;
std::string info_file;
std::string warn_file;
std::string error_file;

// Unified Perfetto timeline (timeline.pftrace)
bool timeline_enabled = true;
std::string timeline_file = "timeline.pftrace";
bool timeline_counters = true;
bool timeline_compress = true;

// Reorder buffer
bool enable_reordering = true;
uint64_t reorder_watermark_cycles = 1000;
size_t reorder_max_events = 100000;

// Simulation metadata
std::string simulation_name;

size_t service_buffer_bytes = 16 * 1024 * 1024;
};

Output Routing​

Log channels (debug/info/warn/error) are text-only and write to events.log. Structured timeline events and lanes go to timeline.pftrace.

Backpressure​

Log and trace queues use the configured policy in both Debug and Release builds:

  • drop: reject the record immediately when the producer queue is full.
  • bounded_wait: publish pending writes, notify the backend and assist bounded ingress while retrying up to backpressure_max_spins (default 4096), then drop if space is still unavailable. This is the default policy.
  • spin_wait: keep retrying while the backend is available, assisting bounded ingress and yielding during prolonged waits. A slow sink can stall the producer.

Assistance uses the same scheduler registration as regular polling and never runs file I/O inside a model callback. Waiting producers stop retrying if the backend stops or fails; records larger than queue capacity are rejected. Therefore spin_wait preserves records under pressure only while the backend can make progress and each record fits its queue.

The consumer publishes freed queue space after each bounded drain, so producers can reuse it without waiting for a larger read-commit batch.

Emergency Flush on Crash​

When a simulation throws a C++ exception, buffered observer data would normally be lost. CrashHandler::emergencyFlush() provides best-effort recovery:

// Called automatically by SimulationApp exception handlers
chronon::sender::CrashHandler::emergencyFlush();

This performs two steps:

  1. ThreadContextManager::flushAll() - commits per-thread SPSC queue write pointers so the backend can see all buffered events
  2. ObservationManager::stopBackend() - drains queues and flushes output files

For fatal signals (SIGSEGV/SIGABRT/etc), the installed signal handler is strict async-signal-safe and exits immediately after printing crash info.

When using SimulationApp, exception-path flushing is handled automatically. For manual setups, install signal handlers and call emergencyFlush() in your own catch blocks:

chronon::sender::CrashHandler::install(); // Signal handlers
try {
sim.runUntilTermination(max_cycles);
} catch (...) {
chronon::sender::CrashHandler::emergencyFlush();
throw;
}