Skip to content

Report audit instance environment, storage footprint and ingestion health in the usage report - #5965

Open
johnsimons wants to merge 9 commits into
masterfrom
john/telemetry_5
Open

johnsimons wants to merge 9 commits into
masterfrom
john/telemetry_5

Conversation

@johnsimons

Copy link
Copy Markdown
Member

Summary

This PR adds facts to the usage report's EnvironmentData about how customers run their audit instances and how much load ingestion puts on each side. The audit instance gets a new GET /api/environment endpoint. The primary reads it from every audit remote once a day and reports the results under Audit.*, combined across instances. Both sides also report their storage footprint, the database edition or service tier, and ingestion rates and health since the process started.

Why

A proposal to ingest audit messages in the primary instance was withdrawn in September 2026, largely because there was no data on how customers run audit instances. A key only appears in reports from the release that adds it, so the next proposal needs this data to already exist when it is written. The EF Core audit persisters added in #5936 may also share the primary's database, so the report should detect that sharing, measure the load on each side, and show whether ingestion keeps up.

What the report gains

Primary instance:

Key Values Notes
Storage.ServerEdition Express, Standard, Enterprise, Other, NotApplicable, Unknown SQL Server edition family. Developer and Evaluation report Enterprise. NotApplicable on Azure SQL Database, Managed Instance, PostgreSQL and RavenDB
Storage.ServiceObjective Basic, Standard, Premium, GeneralPurpose, BusinessCritical, Hyperscale, ElasticPool, Other, NotApplicable, Unknown Azure SQL Database tier, NotApplicable everywhere else
Storage.SizeGB number or Unknown From catalog statistics for the configured schema, so no rows are scanned
Storage.MessageCount count or Unknown Row estimate of the failed messages table
Storage.UnresolvedFailedMessages count
Health.Error.FailedImports count
Health.Error.RetentionBehindHours whole hours How far the oldest resolved or archived message is past the retention window. EF Core only, because RavenDB expiry is engine driven
Health.Error.UptimeHours whole hours
Ingestion.Error.AvgDailyMessages count Messages per day since the process started
Ingestion.Error.BusyPercent 0 to 100 Share of uptime spent processing batches
Ingestion.Error.StorageMsPerMessage number Storage write time per message
Health.Error.LagOver1MinPercent, LagOver10MinPercent, LagOver60MinPercent 0 to 100 Share of messages ingested more than 1, 10 or 60 minutes after their TimeOfFailure header. The audit side uses ProcessingEnded

Audit instances, combined:

Key Combined as
Audit.Host.*, the same eight keys as the primary's Host.* Processor count and memory take the largest value, other keys report Mixed when instances differ
Audit.Storage.Type, RavenServer, Hosting, HostingSource, ServerVersion, ServerEdition, ServiceObjective, FullTextSearch Mixed when instances differ
Audit.Storage.SizeGB, Audit.Storage.MessageCount, Audit.Health.FailedImports Summed, with instances that write to the same store counted once
Audit.Ingestion.AvgDailyMessages Summed across all instances
Audit.Ingestion.BusyPercent and the three Audit.Health.LagOver*Percent keys Largest value
Audit.Health.UptimeHours Smallest value
Audit.SameMachine True, False, Mixed, NotApplicable (containers) or Unknown
Audit.DatabaseSharing SameSchema, SameDatabase, SameServer, SeparateServer, NotApplicable or Unknown, reporting the most shared class across instances

Rates and percentages are left out while a process has been up for less than an hour. Lag keys are left out until a message carries a usable timestamp.

How it works

  • Each audit persister registers environment data providers. RavenDB and the EF Core SQL Server and PostgreSQL persisters are covered. The EF Core probes are copies of the primary's, following the copy-not-share rule for the audit persistence stack.
  • The audit instance serves the values on GET /api/environment, together with a hash of the machine name and hashes of the storage identity: server, database and schema.
  • The primary fetches that endpoint from every remote during its daily audit verification pass. It compares the hashes with its own and stores the per-instance data in the licensing store (RavenDB, EF Core and InMemory). Only the comparison result reaches the report.
  • The primary combines the stored data when the report is built. Only keys on an explicit allowlist are emitted, so a key served by a newer or older audit instance cannot reach the signed report unreviewed.
  • The storage identity resolves an unset schema from the connection (SCHEMA_NAME() on SQL Server, current_schema() on PostgreSQL), so the two sides compare the schema the tables actually live in. A loopback RavenDB address is qualified with the machine name, because localhost names a different server on every machine.
  • On SQL Server the size includes the full-text index, which lives in internal tables. On PostgreSQL the row estimate is summed over the daily partitions of the audit messages table, because the partitioned parent has no estimate of its own.

Privacy and security

These values follow the rules in #5944. Every value is a fixed enum member, a count, a number or a version. The report carries no host names, database names or hashes of them, and no security configuration.

  • The raw SQL Server edition text is not reported. Document the usage report contents and add a coverage decision #5944 dropped the edition from SqlVersion because no analysis needs it. The family keeps the one distinction that matters for capacity, Express and its database size limit. Developer and Evaluation report Enterprise, so the report never shows which licence an instance runs under.
  • The raw Azure service objective is not reported. It encodes the hardware series, and a DC-series name suggests Always Encrypted with secure enclaves, which is encryption configuration.
  • GET /api/environment is anonymous, like the existing GET /api/configuration, because the primary polls it without a user token. The hashes it serves are unsalted SHA-256, so anyone who can reach the audit instance can confirm a guessed machine name or database server name. Where hosts are named after their IP address, as default EC2 Linux host names are, the private IP can be recovered from machine_name_hash by trying every address in the subnet. We accepted this because GET /api/configuration already serves the instance name, log path and queue names in clear text to the same callers. Keyed hashes would not close this gap, because a caller can still test guesses with a key of its own choosing.

Also fixed

ThroughputCollector built each report's EnvironmentData inside the BrokerMetadata.Data dictionary returned by the licensing store. When no broker metadata is stored, as with the Learning transport, the RavenDB and EF Core stores return a shared static default. Every report therefore wrote into the same dictionary. Keys from one report leaked into the next, and two reports built at the same time could corrupt it. Acceptance tests run in parallel in one process, so the two Licensing tests shared that dictionary. On this branch, where building the report takes longer, 3 of 6 SQL Server runs failed with a duplicate key, a timeout, or a value from the other test. The report now copies the dictionary. The bug exists on master as well.

Alternatives considered

  • An hourly collector that kept 400 days of ingestion history in the licensing stores. It was replaced by averages since process start. Volume over time can already be derived from the per-endpoint daily counts the report carries. The averages cover everything else the history measured except the busiest hour, and the history needed new members on all three licensing stores.
  • Taking audit facts only from GET /api/configuration, the channel the coverage decision uses. That response has no host, storage or ingestion facts. The decision itself names a dedicated audit endpoint as the channel if these facts are needed.
  • Hashing /etc/machine-id on Linux. The machine name is used on every OS instead. Generic host names such as localhost can make two machines compare equal. Machine-id has the mirror problem, because cloned VMs share it.
  • Passing every key an audit instance serves straight into the report. The allowlist makes the path fail closed instead.

Known gaps

  • Workers started with --error-ingestion-only are not counted.
  • Audit instances from earlier releases do not serve the endpoint, so the report leaves their keys out.
  • Audit values can be up to a day old, because they come from the daily verification pass.
  • Averages reset when a process restarts. UptimeHours shows how much time they cover.
  • Storage.SizeGB covers the whole schema. When the primary and an audit instance share a schema, both report the same size, and Audit.DatabaseSharing reads SameSchema.
  • PostgreSQL 11 has no pg_partition_tree, so the audit footprint reports Unknown there.

The audit instance serves host and storage facts on a new anonymous GET /api/environment endpoint, gathered from audit-local environment data providers. The primary polls every remote alongside the daily audit verification pass through the new IConfigurationApi.GetRemoteEnvironments, stores the per-instance dictionaries as AuditEnvironmentMetadata in the licensing stores, and the usage report emits them under the Audit.* prefix aggregated across instances: agreement passes through, processor count and memory take the maximum, and disagreement reports Mixed. Remotes that predate the endpoint are omitted rather than guessed.
…or server

Each instance hashes what identifies the machine it runs on and the storage it writes to: the OS machine id on Linux, the machine name elsewhere, NotApplicable in containers, and per engine the RavenDB server URL and database, SQL Server's ServerName and database, or PostgreSQL's cluster system identifier and database, plus the schema. The audit instance serves the hashes on the environment endpoint, the primary compares them with its own during collection, and only the classification reaches the report: Audit.SameMachine as True, False, Mixed, NotApplicable or Unknown, and Audit.DatabaseSharing as SameSchema, SameDatabase, SameServer, SeparateServer, NotApplicable or Unknown, aggregated to the most shared class across instances. A side that cannot be identified reports Unknown rather than guessing from configuration the two sides may spell differently.
…d audit instances

Storage.SizeGB, Storage.MessageCount and Storage.UnresolvedFailedMessages come from catalog statistics scoped to the configured schema on SQL Server and PostgreSQL and from database and collection statistics on RavenDB, never from counting rows. Storage.ServerEdition carries SQL Server's own edition designation and Storage.ServiceObjective the Azure SQL service objective, both NotApplicable elsewhere. The audit instance serves the same keys over the environment endpoint, and the report sums sizes and counts across audit instances with instances sharing a database counted once, keyed by the identity hashes from the sharing detection. A catalog that has no answer, such as a row estimate on a never analysed PostgreSQL table, reports Unknown rather than a guess.
…sides

Both ingestion pipelines keep running totals in memory: messages stored per completed batch, time spent processing batches, storage write time on the error side, and how far behind the endpoint's own timestamp each message arrived, bucketed at one, ten and sixty minutes. The audit instance serves its counters on the environment endpoint. A new hourly collector in the licensing component polls the primary's counters in process and every audit instance's over the API, folds the deltas into daily records in the licensing store with the audit side summed across instances, and keeps 400 days. A restarted process contributes its whole counters as the delta, a counter set seen for the first time only establishes a baseline, and a collector restart builds on what today's record already holds. The report emits per side the average daily messages, the busiest hour's messages and its busy share, plus storage milliseconds per message for the error side, from the days inside the report window.

Known gap, accepted in the plan: --error-ingestion-only workers keep their own counters and nothing polls them.
Health.Error.FailedImports and Health.Error.RetentionBehindHours come from the primary's persister: failed import rows counted from the store, and how many hours the oldest failed message sits past the retention window on the EF persisters, with RavenDB omitting the retention key because expiry is engine driven. The audit instance serves its failed audit import count on the environment endpoint, summed across instances with shared databases counted once. Restarts are the process start time changes the hourly collector observes between polls, and the lag buckets it accumulates become Health.*.LagOver1MinPercent, LagOver10MinPercent and LagOver60MinPercent over the report window, emitted only when any message carried a usable timestamp.
…on the environment endpoint

Instead of polling ingestion counters hourly, accumulating deltas into daily records in the licensing store, and reading those records back into the report, each instance now computes ingestion and health summaries directly from its running totals at report time.

Both the audit and error sides expose their counters through IEnvironmentDataProvider implementations that call IngestionSummary.Describe, which yields keys like Ingestion.AvgDailyMessages, Ingestion.BusyPercent, Health.UptimeHours, and Health.LagOver*Percent relative to the process uptime. The audit instance serves these on the environment endpoint alongside its existing data; the primary aggregates them under the Audit.* prefix when building the report, summing daily message rates across instances and taking the worst saturation and lag values.

This removes IngestionHistoryCollectorHostedService, IngestionHistory and IngestionDay contracts, IErrorIngestionSnapshotProvider, GetRemoteIngestionCounters from IConfigurationApi, and GetIngestionHistory/SaveIngestionHistory from ILicensingDataStore and both persistence backends.
…sisters

SQL Server and PostgreSQL audit persistence now probe how their database is hosted, how much it stores, and what identifies it, the same way the primary instance's EF persisters already did, so the report gets real Storage.* data for audit instances on these backends instead of nothing. SQL Server's edition and service objective are normalised to a small fixed set of categories instead of raw server strings, and both EF persisters resolve the configured schema the same way when none is set.

AuditEnvironmentDataAggregator now passes through only an explicit whitelist of known keys, so a key served by an audit instance of another version can never reach the report unreviewed. ThroughputCollector copies broker metadata into a new dictionary before it reaches the report, so the stored metadata is never mutated. The environment endpoint's machine hash field is renamed from machine_id_hash to machine_name_hash to match what it actually hashes, and RavenDB's storage identity now treats a loopback address as this machine rather than a shared one.
@johnsimons johnsimons self-assigned this Oct 7, 2026
…t test

WaitForFullTextIndex is now exposed as a public method with a default cancellation token so the environment data test can reuse the same wait the search tests rely on, instead of racing the index population. The test now waits for the full-text index to catch up before probing storage footprint, and asserts the footprint is only greater than the tables size rather than an exact sum, since the full-text index size can still drift slightly after the wait. The duplicated inline SQL for computing table and full-text sizes is consolidated into a single SizeGB helper.
…ition

Catalog reads under READ COMMITTED can deadlock with concurrent DDL, so the probe now runs under READ UNCOMMITTED instead. That isolation level can let a scan lose its place when pages move, so both error codes are now retried up to three times with a short backoff before the probe gives up and returns null. Applied to both the primary and audit EF Core SQL Server persisters, and the audit environment data test's inline SQL is updated to match the new isolation level.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant