Metrics and monitoring
This page describes the current OSA Proxy implementation. The archived Java/Spring implementation is available in Archived Java/Spring implementation.
OSA Proxy exposes Prometheus-format metrics at:
Example:
The Go implementation does not provide Spring Boot actuator endpoints. Scrape
/metrics, not /actuator/metrics or /actuator/prometheus.
What OSA Proxy provides
OSA Proxy provides:
- OSA Proxy application metrics and standard Go/process metrics;
- example Prometheus, Alertmanager, and Docker Compose configurations in the OSA Proxy repository.
Before using recording rules or alerts in production, adjust thresholds, intervals, Alertmanager routing, and receivers to match your load, SLOs, and incident response procedures. CodeScoring cannot define universal CPU, memory, latency, or no-traffic thresholds for every installation.
Connecting Prometheus
OSA Proxy does not push metrics. The monitoring system must scrape /metrics.
Metric transport and storage configuration are outside the scope of OSA Proxy.
For queries and rules, use the service="osa-proxy" and environment
target labels. Prometheus adds an instance label for every target, which
distinguishes OSA Proxy replicas.
OSA Proxy exports a constant osa_proxy_info metric with the value 1. It
selects only OSA Proxy targets and prevents identically named go_*,
process_*, and HTTP metrics from other services from being mixed in.
Prometheus also adds the technical job label automatically. Its value depends
on target discovery and is not part of the OSA Proxy contract; queries and
rules should not require a specific job value.
Static targets
If one Prometheus server scrapes multiple instances, use one target group.
Prometheus assigns a different instance to every address:
The resulting series differ by instance:
Do not expose /metrics to the internet. The endpoint is intended for an
internal monitoring system.
HTTP outcome semantics
User traffic and operational probes are separated:
/healthz,/readyz, and/metricsare recorded inhttp_operational_requests_totaland excluded from user-traffic SLIs;outcome="success"represents a successful user request;outcome="blocked"represents a policy decision and is not a technical service failure;outcome="client_error"represents an invalid client request;outcome="error"represents an OSA Proxy or dependency failure.
For upstream registries, an expected HTTP 404 has the not_found outcome,
other 4xx responses use client_error, and transport errors or 5xx
responses use error. Build error ratios from outcome="error", not from all
4xx and 5xx responses combined.
Application metric contract
The tables below list the OSA Proxy-specific metrics. The separately
registered promhttp_metric_handler_errors_total and the standard Go/process
collector families are described under Runtime and metrics endpoint.
For a histogram, Prometheus automatically creates series with the _bucket,
_sum, and _count suffixes; they are parts of one metric and are not listed
separately. A counter whose name ends in _total is exposed under that same
name.
HTTP, CodeScoring, and upstream
Scanner, manifests, and cache
Artifactory and Nexus discovery
Each provider exports the same metric groups:
Managed configuration
Revision gauges show whether a replica has activated the desired state;
managed_configuration_convergence_lag should return to 0. Readiness and
health gauges use 1 for healthy/available and 0 otherwise. Store state is a
one-hot enum with disabled, healthy, degraded, and never_loaded values.
Durability state is a one-hot enum with disabled, verified, unverified,
and unsafe_accepted values. Activation labels use stage="build|activation",
result="success|failure", and failure="none|build|start|activation".
Discovery health uses provider="artifactory|nexus".
Lease and managed-discovery health are relevant only when Configuration Store is enabled. The bundled baseline rules do not currently define alerts for these managed-configuration series; add deployment-specific rules when they are required by the operational policy.
Runtime and metrics endpoint
The standard Go collectors provide go_* and process_* metrics for CPU, RSS,
heap, GC, goroutines, and file descriptors.
promhttp_metric_handler_errors_total{cause} is the standard counter of
/metrics gathering or encoding errors.
OSA Proxy does not export network byte counts or network-interface utilization.
In Kubernetes, use kubelet/cAdvisor metrics such as
container_network_receive_bytes_total and
container_network_transmit_bytes_total, filtered by namespace and pod.
When series may be absent
A missing series does not always indicate a failure:
- whitelist mode does not generate CodeScoring API, retry, circuit-breaker, or fallback traffic;
- cache metrics appear when Redis verdict caching is enabled and operations occur;
- proactive refresh metrics require the background task to be enabled;
- Artifactory and Nexus discovery metrics require the provider to be enabled and at least one sync attempt;
- histogram series appear after the first matching observation.
Do not replace missing latency with zero because zero latency represents a completed instantaneous request. Return a zero error ratio only when traffic exists but no errors occurred.
PromQL examples
User request rate by replica and outcome:
Inbound technical error ratio:
Successful request p95 latency:
Age of the last successful Artifactory discovery:
Recording-rule metrics
When recording rules are loaded, Prometheus calculates seven additional series.
These are not application metrics: they are absent from /metrics and exist
only after the recording rules have been loaded.
Baseline alerts
Recording rules and baseline alerts can cover the main failure classes:
Adapt the example thresholds to the actual load. Disable or adjust the no-traffic alert for environments where idle periods are expected.
Pay particular attention to codescoring_fallbacks_total{result="allow"}: even
a single fail-open decision may be significant. Detect it with
increase(...[5m]) > 0, rather than only a long-running rate() > 0 condition
with a large for duration.
Exemplars and tracing
When OpenTelemetry tracing is enabled, histogram metrics may include exemplars
with a trace_id. If Prometheus and Grafana are configured to store and display
exemplars, operators can navigate from a latency point to the corresponding
trace.
Never add trace_id as a regular metric label because that creates a separate
series for every request.
Post-installation validation
-
Check liveness and readiness:
-
Check the metrics endpoint:
-
Check the identity metric:
-
Open the Prometheus Targets page and verify that OSA Proxy is
UP. -
Run the
osa_proxy_infoquery and verify theservice,environment, andinstancelabels. The technicaljoblabel is also present, but its specific value is not part of the OSA Proxy configuration contract. -
Verify that Prometheus loaded the recording/alerting rules and that Alertmanager accepts a test alert.
