Mediator Details
The mediator is a lightweight container you run inside your own environment. It turns your telemetry into a live, abstract representation of your system (entities, their relationships and their symptom states), and sends only that representation to Causely's reasoning system. Your raw data stays with you. That design is why Causely works at large-scale telemetry volumes.
- It processes telemetry where it lives. Metrics, traces, logs, events and alerts are analyzed inside your own cluster or VPC.
- It sends an abstract representation, not your data. What goes upstream is a compact, compressed snapshot of entities, their relationships and their symptom states (healthy or anomalous). That's the input causal reasoning needs.
- It scales out, not just up. Heavy trace volumes can be spread across horizontally scaled trace workers. Large estates run many mediators, one per cluster, VPC or business domain. The reasoning system joins them into one causal graph.
- It works with the telemetry you already have. The mediator combines data from all of your existing sources into one picture. You never have to choose one source of truth.
For how the mediator fits with the other components Causely deploys, see Architecture.
Where the mediator sits
The mediator runs next to your workloads: one per Kubernetes cluster, ECS or Nomad environment, or VM estate. It is the only Causely component that processes your telemetry. It connects to the Causely backend over a single outbound TLS connection. That backend can be Causely-hosted or run in your own cloud (bring your own cloud, BYOC).
Telemetry arrives in two ways: sources push it to the mediator, or the mediator pulls it.
- Pushed to the mediator: OpenTelemetry traces and metrics (OTLP on port 4317), Datadog APM traffic, and alert webhooks.
- Pulled by the mediator: it queries your existing tools and platform APIs on a schedule (Prometheus, Datadog, Dynatrace, cloud provider APIs, databases, message brokers, CMDB and inventory APIs). Nothing needs to be re-instrumented. See Supported Telemetry for the full list.
- Local store: entity metrics are written to a small time series store in your environment, kept for one day by default. You can point this at your own Prometheus-compatible store instead.
What the mediator does, continuously
The mediator is always on. It keeps its representation of your system current in real time. It runs six steps in a loop.
-
Discover. It builds a live graph of your entities (services, workloads, pods, nodes, databases, queues, topics, HTTP routes, cloud resources) and their relationships. It uses Kubernetes and cloud APIs, your APM and monitoring systems, and CMDB or inventory data.
-
Ingest. It receives OpenTelemetry traces and metrics, scrapes your existing tools on a schedule (as often as every 10 seconds), and scans container logs on a schedule for errors and known failure patterns.
-
Distill. Each span is reduced as soon as it arrives:
- Request, error and duration counters, plus compact latency sketches.
- URL paths collapsed into route templates, so
/users/123becomes/users/{id}. - SQL statements normalized, with literal values replaced by placeholders.
Once distilled, the raw span is discarded.
-
Detect locally. Symptoms such as high latency, error spikes, saturation, restarts and consumer lag are evaluated inside your environment. They use default thresholds, your overrides, or thresholds the mediator learns from your own recent history. Learned thresholds are recalculated hourly from the last 24 hours.
-
Sync. Approximately every 30 seconds, the mediator sends its current representation to the reasoning system: entities, relationships and symptom states. The snapshot is compressed and streamed in chunks.
-
Answer. When the reasoning system is building or explaining a Diagnosis, it can ask for specific evidence. Examples are one metric series for one entity, or a short, scrubbed log excerpt for one pod. The mediator fetches just that item and returns it.
This is how Causely can reason over your entire system continuously without first copying all of your data to a central store.
What happens to each type of telemetry
Each type of telemetry becomes structure or state in the mediator's abstract representation, and is then discarded or kept locally. Only that representation leaves your environment continuously. Anything more is sent only when the reasoning system asks for a specific piece of evidence.
| Telemetry | Where it comes from | What the mediator does with it | What stays in your environment | What goes to the reasoning system |
|---|---|---|---|---|
| Traces | OpenTelemetry SDKs and collectors, Datadog APM and other APM agents | Discovers dependencies (HTTP, gRPC, database, queue and topic, GenAI and MCP calls). Computes request rate, error rate and latency per service and endpoint. | Nothing. Spans are discarded once analyzed. | Relationships, plus symptom states such as high latency or errors |
| Metrics | Prometheus, Datadog, Dynatrace, cloud monitoring, OTLP metrics, Causely agents | Maps them onto the entities they describe. Evaluates symptoms against thresholds. | Entity metrics in the local store (one-day default retention) | Symptom states. A specific series only on request. |
| Logs | Container logs, read through the Kubernetes API or from Loki, Elasticsearch or Splunk | Scans new lines on a schedule. Counts error lines and known failure patterns (out-of-memory, deadlocks, connection-pool exhaustion, and similar) and turns them into symptoms. | Your log store stays the system of record. Causely keeps only counts. | Symptom states. A short, PII-scrubbed excerpt only on request. |
| Events and alerts | Kubernetes events, Alertmanager, Prometheus, Grafana, Datadog, Dynatrace, incident management systems | Maps them to symptoms on the affected entity. You can map your own alerts. | Nothing extra | Alert and symptom states |
| Topology and config | Kubernetes, AWS, Azure, GCP, service mesh, APM topology, CMDB, topology files | Builds and maintains the entity graph. Merges overlapping sources. | Configuration details | Entities and relationships. Config only on request. |
Logs
Causely doesn't copy, index, or store your logs. The mediator queries them where they already live and pulls only what it needs: it retrieves the relevant logs as supporting evidence when it identifies an active issue and scans on a schedule for defined error patterns (such as Java null pointer exceptions) that become symptoms for Causely's inference system.
- Scheduled, incremental scanning. The mediator works through every container on a rolling cycle, reading only the lines written since its last pass. It reads from the Kubernetes API, or from Loki, Elasticsearch or Splunk if you use them. Scans run with bounded concurrency, so your log store isn't hit by a constant stream of heavy queries.
- Patterns become symptoms. Each pass counts error lines and matches a built-in library of known failure signatures, such as out-of-memory errors, deadlocks, connection-pool exhaustion, full disks, open circuit breakers and thread-pool exhaustion. You can add your own patterns. Matches activate symptoms on the affected entity, alongside symptoms from metrics and traces.
- Only counts are kept. Log lines are discarded once they're counted. Nothing is indexed, and no log stream goes upstream.
- Evidence on demand, scrubbed. When the reasoning system needs to show evidence for a specific entity, the mediator retrieves a short, scoped slice of recent lines. It removes emails, phone numbers, card numbers, national IDs, IP addresses, tokens, API keys and credentials before anything leaves.
Today, service dependencies come from traces, the service mesh, and cloud and APM topology rather than from log text, because those sources record who calls whom directly. If your logs carry signals those sources miss, custom patterns are the simplest way to bring them in.
The result is that Causely uses the signal in your logs without you needing to move them or depend on storing and indexing them, however large your log volume is.
Combine data from multiple sources
You don't have to choose one source of truth. The mediator merges every source it receives into a single entity graph, so each source adds to the same picture instead of producing its own partial one.
For example, one team instruments its services with OpenTelemetry and another uses Datadog APM. Both send traces to the mediator. It recognizes the same services and endpoints in each, so a call from an OpenTelemetry-instrumented service to a Datadog-instrumented one becomes one dependency in one graph, not two disconnected maps.
- Same thing, same entity. Entities get stable IDs derived from what they are, such as their Kubernetes identity, cloud resource or service name. When several sources describe the same service, they update one entity, not several.
- Conflicts resolve per attribute. When two sources disagree on a value, the mediator applies a defined source priority for that type of data. Each attribute comes from the highest-priority source available, and gaps are filled from the rest.
- Nothing to rip out. Keep your collectors, vendors and dashboards. Causely adds a causal layer on top of them.
Where a service has no instrumentation at all, Causely's automatic instrumentation can add coverage without code changes, and its data merges the same way.
Telemetry from a service must reach the same mediator that discovers it. The mediator places incoming traces into the entity graph by matching them to workloads it already knows about, so a trace from a service it can't see has nowhere to go. That's why mediators run close to the workloads they observe, not as a single central sink at the network edge.
How it scales
The mediator's work grows with the number of entities it manages, not with the volume of telemetry flowing through it. The answer to "can it handle petabytes a day?" therefore doesn't depend on any single component. Volume is absorbed where the data is produced, and capacity is added by adding trace workers and mediators.
Every request to the same endpoint updates the same counters and the same latency sketch. Ten requests a second and ten thousand a second use roughly the same memory. What adds memory is new entities: services, pods, routes, queues, database queries. A cluster with more services needs more mediator memory; a high-traffic cluster with the same number of services does not.
The mediator works from a statistically significant sample of traces and metrics. It does not need 100% of raw telemetry to perform accurate analysis. When incoming trace volume exceeds what the mediator can handle under memory pressure, it drops the excess before decoding it and keeps processing. This is intentional, not data loss: it means the mediator already has enough data for analysis, and symptom detection stays accurate. It doesn't fall over, and it doesn't back-pressure your collectors. Request counts you see in Causely may therefore differ from totals in other monitoring systems.
Three ways to add capacity
- Scale up the mediator. It starts small and grows with the environment. See Recommended Sizing Guidelines for typical requests and limits, and Sizing the Mediator for large environments to raise the default 16Gi memory limit.
- Scale out trace processing. For high trace volumes, turn on the scalable trace processor. A stateless trace controller receives traces and routes each service to a dedicated trace worker. Workers run as a horizontally scaled set and do the CPU-heavy span analysis. The mediator only receives their results. Routing can use
service.nameor a coarser key, such as a business application, to keep related services together. - Run more mediators. Large estates deploy one mediator per cluster, VPC or business domain. Each mediator owns a subset of services. The reasoning system joins their entity graphs into one causal graph, including dependencies that cross mediator boundaries. If your traces flow through a central OpenTelemetry Collector gateway, its load-balancing exporter can route each service's traces to the right mediator.
Egress at scale
What streams upstream continuously is an abstract representation of your environment: entity IDs, relationships and symptom states, compressed and sent approximately every 30 seconds. Its size depends on how many entities you have, not how much telemetry they produce, and growth in data volume does not become growth in egress.
Why run the mediator in your environment
Running the mediator inside your environment means the analysis happens where the data lives, which pays off in five ways:
- Egress stays small. Only the abstract representation and targeted evidence leave, not 100% of traces, metrics and logs.
- The footprint stays light. The mediator focuses on real-time analysis and keeps no long-term history.
- Your data stays local. Raw telemetry remains in your datacenter, and mediators can run at the edge of your infrastructure, so data doesn't cross regions or clouds.
- The picture is always current. Topology and symptom states are up to date before an incident begins.
- Detection sees all traffic. Symptoms are evaluated in the mediator, before any sampling your pipeline applies downstream.
It also means Causely runs continuously inside your environment, rather than as a destination you ship data to. Three familiar approaches show what that avoids:
- Sending all telemetry to a central store. Ingest, indexing and cloud egress costs rise with every new service and every traffic spike. Teams sample aggressively to control that cost, and the sampled-away requests become blind spots. Raw telemetry also leaves your environment, which triggers data-sovereignty and compliance reviews. What you get is stored data, not a causal model.
- Querying your tools when an incident starts. Topology and dependencies are rebuilt, or guessed, at the start of every investigation. Each investigation makes many API and query calls, which adds latency, runs into rate limits and costs money in query fees and LLM tokens. It only sees what your tools kept after sampling and retention limits.
- Building on a data lake. You pay to store everything to analyze a small fraction of it. Analysis happens at query time, which is slowest when an incident is in progress. Correlation rules need constant upkeep and still produce correlation, not causation.
The outcome is causal answers from a live picture of your entire system, built where your data already lives.
Data handling
Raw telemetry stays in your environment. The abstract representation leaves continuously, and targeted, scrubbed evidence leaves only when the reasoning system asks for it. The mediator connects outbound over TLS only; Causely never opens inbound connections into your environment. For the full breakdown of what is sent, what is never sent and how personal data is scrubbed, see What leaves your environment.
Frequently asked questions
Do we need to send Causely all of our logs? No. The mediator scans your logs where they live and keeps only counts of errors and failure patterns, which become symptoms. Short, scrubbed excerpts are retrieved only as evidence for a specific entity.
What happens if we send more traces than a mediator can process? The mediator sheds the excess before decoding it and keeps running on a statistically significant sample. If that happens regularly, turn on the scalable trace processor or split the workload across more mediators.
Do we need one mediator for the whole company? No, and we recommend against it. Deploy one mediator per cluster, VPC or business domain. The reasoning system combines them into one causal graph.
Do we have to replace our collectors or observability vendors? No. Keep them. The mediator pulls from the tools you have and accepts standard OpenTelemetry, and it merges every source into one entity graph.
How much network egress should we expect? It depends on the size of your topology, not your data volume. The continuous stream is a compressed snapshot of entities, relationships and symptom states approximately every 30 seconds, plus occasional evidence requests.
Can we run everything on our own infrastructure? Yes. The reasoning system can run in your own cloud (bring your own cloud, BYOC). With Causely's own health telemetry turned off, nothing leaves your network.