Repository navigation
Report audit instance environment, storage footprint and ingestion health in the usage report - #5965
Open
johnsimons wants to merge 9 commits into
Open
johnsimons wants to merge 9 commits into
johnsimons wants to merge 9 commits into
Conversation
The audit instance serves host and storage facts on a new anonymous GET /api/environment endpoint, gathered from audit-local environment data providers. The primary polls every remote alongside the daily audit verification pass through the new IConfigurationApi.GetRemoteEnvironments, stores the per-instance dictionaries as AuditEnvironmentMetadata in the licensing stores, and the usage report emits them under the Audit.* prefix aggregated across instances: agreement passes through, processor count and memory take the maximum, and disagreement reports Mixed. Remotes that predate the endpoint are omitted rather than guessed.
…or server Each instance hashes what identifies the machine it runs on and the storage it writes to: the OS machine id on Linux, the machine name elsewhere, NotApplicable in containers, and per engine the RavenDB server URL and database, SQL Server's ServerName and database, or PostgreSQL's cluster system identifier and database, plus the schema. The audit instance serves the hashes on the environment endpoint, the primary compares them with its own during collection, and only the classification reaches the report: Audit.SameMachine as True, False, Mixed, NotApplicable or Unknown, and Audit.DatabaseSharing as SameSchema, SameDatabase, SameServer, SeparateServer, NotApplicable or Unknown, aggregated to the most shared class across instances. A side that cannot be identified reports Unknown rather than guessing from configuration the two sides may spell differently.
…d audit instances Storage.SizeGB, Storage.MessageCount and Storage.UnresolvedFailedMessages come from catalog statistics scoped to the configured schema on SQL Server and PostgreSQL and from database and collection statistics on RavenDB, never from counting rows. Storage.ServerEdition carries SQL Server's own edition designation and Storage.ServiceObjective the Azure SQL service objective, both NotApplicable elsewhere. The audit instance serves the same keys over the environment endpoint, and the report sums sizes and counts across audit instances with instances sharing a database counted once, keyed by the identity hashes from the sharing detection. A catalog that has no answer, such as a row estimate on a never analysed PostgreSQL table, reports Unknown rather than a guess.
…sides Both ingestion pipelines keep running totals in memory: messages stored per completed batch, time spent processing batches, storage write time on the error side, and how far behind the endpoint's own timestamp each message arrived, bucketed at one, ten and sixty minutes. The audit instance serves its counters on the environment endpoint. A new hourly collector in the licensing component polls the primary's counters in process and every audit instance's over the API, folds the deltas into daily records in the licensing store with the audit side summed across instances, and keeps 400 days. A restarted process contributes its whole counters as the delta, a counter set seen for the first time only establishes a baseline, and a collector restart builds on what today's record already holds. The report emits per side the average daily messages, the busiest hour's messages and its busy share, plus storage milliseconds per message for the error side, from the days inside the report window. Known gap, accepted in the plan: --error-ingestion-only workers keep their own counters and nothing polls them.
Health.Error.FailedImports and Health.Error.RetentionBehindHours come from the primary's persister: failed import rows counted from the store, and how many hours the oldest failed message sits past the retention window on the EF persisters, with RavenDB omitting the retention key because expiry is engine driven. The audit instance serves its failed audit import count on the environment endpoint, summed across instances with shared databases counted once. Restarts are the process start time changes the hourly collector observes between polls, and the lag buckets it accumulates become Health.*.LagOver1MinPercent, LagOver10MinPercent and LagOver60MinPercent over the report window, emitted only when any message carried a usable timestamp.
…on the environment endpoint Instead of polling ingestion counters hourly, accumulating deltas into daily records in the licensing store, and reading those records back into the report, each instance now computes ingestion and health summaries directly from its running totals at report time. Both the audit and error sides expose their counters through IEnvironmentDataProvider implementations that call IngestionSummary.Describe, which yields keys like Ingestion.AvgDailyMessages, Ingestion.BusyPercent, Health.UptimeHours, and Health.LagOver*Percent relative to the process uptime. The audit instance serves these on the environment endpoint alongside its existing data; the primary aggregates them under the Audit.* prefix when building the report, summing daily message rates across instances and taking the worst saturation and lag values. This removes IngestionHistoryCollectorHostedService, IngestionHistory and IngestionDay contracts, IErrorIngestionSnapshotProvider, GetRemoteIngestionCounters from IConfigurationApi, and GetIngestionHistory/SaveIngestionHistory from ILicensingDataStore and both persistence backends.
…sisters SQL Server and PostgreSQL audit persistence now probe how their database is hosted, how much it stores, and what identifies it, the same way the primary instance's EF persisters already did, so the report gets real Storage.* data for audit instances on these backends instead of nothing. SQL Server's edition and service objective are normalised to a small fixed set of categories instead of raw server strings, and both EF persisters resolve the configured schema the same way when none is set. AuditEnvironmentDataAggregator now passes through only an explicit whitelist of known keys, so a key served by an audit instance of another version can never reach the report unreviewed. ThroughputCollector copies broker metadata into a new dictionary before it reaches the report, so the stored metadata is never mutated. The environment endpoint's machine hash field is renamed from machine_id_hash to machine_name_hash to match what it actually hashes, and RavenDB's storage identity now treats a loopback address as this machine rather than a shared one.
…t test WaitForFullTextIndex is now exposed as a public method with a default cancellation token so the environment data test can reuse the same wait the search tests rely on, instead of racing the index population. The test now waits for the full-text index to catch up before probing storage footprint, and asserts the footprint is only greater than the tables size rather than an exact sum, since the full-text index size can still drift slightly after the wait. The duplicated inline SQL for computing table and full-text sizes is consolidated into a single SizeGB helper.
…ition Catalog reads under READ COMMITTED can deadlock with concurrent DDL, so the probe now runs under READ UNCOMMITTED instead. That isolation level can let a scan lose its place when pages move, so both error codes are now retried up to three times with a short backoff before the probe gives up and returns null. Applied to both the primary and audit EF Core SQL Server persisters, and the audit environment data test's inline SQL is updated to match the new isolation level.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR adds facts to the usage report's
EnvironmentDataabout how customers run their audit instances and how much load ingestion puts on each side. The audit instance gets a newGET /api/environmentendpoint. The primary reads it from every audit remote once a day and reports the results underAudit.*, combined across instances. Both sides also report their storage footprint, the database edition or service tier, and ingestion rates and health since the process started.Why
A proposal to ingest audit messages in the primary instance was withdrawn in September 2026, largely because there was no data on how customers run audit instances. A key only appears in reports from the release that adds it, so the next proposal needs this data to already exist when it is written. The EF Core audit persisters added in #5936 may also share the primary's database, so the report should detect that sharing, measure the load on each side, and show whether ingestion keeps up.
What the report gains
Primary instance:
Storage.ServerEditionExpress,Standard,Enterprise,Other,NotApplicable,UnknownEnterprise.NotApplicableon Azure SQL Database, Managed Instance, PostgreSQL and RavenDBStorage.ServiceObjectiveBasic,Standard,Premium,GeneralPurpose,BusinessCritical,Hyperscale,ElasticPool,Other,NotApplicable,UnknownNotApplicableeverywhere elseStorage.SizeGBUnknownStorage.MessageCountUnknownStorage.UnresolvedFailedMessagesHealth.Error.FailedImportsHealth.Error.RetentionBehindHoursHealth.Error.UptimeHoursIngestion.Error.AvgDailyMessagesIngestion.Error.BusyPercentIngestion.Error.StorageMsPerMessageHealth.Error.LagOver1MinPercent,LagOver10MinPercent,LagOver60MinPercentTimeOfFailureheader. The audit side usesProcessingEndedAudit instances, combined:
Audit.Host.*, the same eight keys as the primary'sHost.*Mixedwhen instances differAudit.Storage.Type,RavenServer,Hosting,HostingSource,ServerVersion,ServerEdition,ServiceObjective,FullTextSearchMixedwhen instances differAudit.Storage.SizeGB,Audit.Storage.MessageCount,Audit.Health.FailedImportsAudit.Ingestion.AvgDailyMessagesAudit.Ingestion.BusyPercentand the threeAudit.Health.LagOver*PercentkeysAudit.Health.UptimeHoursAudit.SameMachineTrue,False,Mixed,NotApplicable(containers) orUnknownAudit.DatabaseSharingSameSchema,SameDatabase,SameServer,SeparateServer,NotApplicableorUnknown, reporting the most shared class across instancesRates and percentages are left out while a process has been up for less than an hour. Lag keys are left out until a message carries a usable timestamp.
How it works
GET /api/environment, together with a hash of the machine name and hashes of the storage identity: server, database and schema.SCHEMA_NAME()on SQL Server,current_schema()on PostgreSQL), so the two sides compare the schema the tables actually live in. A loopback RavenDB address is qualified with the machine name, becauselocalhostnames a different server on every machine.Privacy and security
These values follow the rules in #5944. Every value is a fixed enum member, a count, a number or a version. The report carries no host names, database names or hashes of them, and no security configuration.
SqlVersionbecause no analysis needs it. The family keeps the one distinction that matters for capacity, Express and its database size limit. Developer and Evaluation reportEnterprise, so the report never shows which licence an instance runs under.GET /api/environmentis anonymous, like the existingGET /api/configuration, because the primary polls it without a user token. The hashes it serves are unsalted SHA-256, so anyone who can reach the audit instance can confirm a guessed machine name or database server name. Where hosts are named after their IP address, as default EC2 Linux host names are, the private IP can be recovered frommachine_name_hashby trying every address in the subnet. We accepted this becauseGET /api/configurationalready serves the instance name, log path and queue names in clear text to the same callers. Keyed hashes would not close this gap, because a caller can still test guesses with a key of its own choosing.Also fixed
ThroughputCollectorbuilt each report'sEnvironmentDatainside theBrokerMetadata.Datadictionary returned by the licensing store. When no broker metadata is stored, as with the Learning transport, the RavenDB and EF Core stores return a shared static default. Every report therefore wrote into the same dictionary. Keys from one report leaked into the next, and two reports built at the same time could corrupt it. Acceptance tests run in parallel in one process, so the two Licensing tests shared that dictionary. On this branch, where building the report takes longer, 3 of 6 SQL Server runs failed with a duplicate key, a timeout, or a value from the other test. The report now copies the dictionary. The bug exists on master as well.Alternatives considered
GET /api/configuration, the channel the coverage decision uses. That response has no host, storage or ingestion facts. The decision itself names a dedicated audit endpoint as the channel if these facts are needed./etc/machine-idon Linux. The machine name is used on every OS instead. Generic host names such aslocalhostcan make two machines compare equal. Machine-id has the mirror problem, because cloned VMs share it.Known gaps
--error-ingestion-onlyare not counted.UptimeHoursshows how much time they cover.Storage.SizeGBcovers the whole schema. When the primary and an audit instance share a schema, both report the same size, andAudit.DatabaseSharingreadsSameSchema.pg_partition_tree, so the audit footprint reportsUnknownthere.