Metrics and monitoring

OSA Proxy implementation

This page describes the current OSA Proxy implementation. The archived Java/Spring implementation is available in Archived Java/Spring implementation.

OSA Proxy exposes Prometheus-format metrics at:

GET /metrics

Example:

curl http://localhost:8080/metrics

The Go implementation does not provide Spring Boot actuator endpoints. Scrape /metrics, not /actuator/metrics or /actuator/prometheus.

What OSA Proxy provides

OSA Proxy provides:

  • OSA Proxy application metrics and standard Go/process metrics;
  • example Prometheus, Alertmanager, and Docker Compose configurations in the OSA Proxy repository.
The baseline configuration is an example

Before using recording rules or alerts in production, adjust thresholds, intervals, Alertmanager routing, and receivers to match your load, SLOs, and incident response procedures. CodeScoring cannot define universal CPU, memory, latency, or no-traffic thresholds for every installation.

Connecting Prometheus

OSA Proxy does not push metrics. The monitoring system must scrape /metrics. Metric transport and storage configuration are outside the scope of OSA Proxy.

For queries and rules, use the service="osa-proxy" and environment target labels. Prometheus adds an instance label for every target, which distinguishes OSA Proxy replicas.

OSA Proxy exports a constant osa_proxy_info metric with the value 1. It selects only OSA Proxy targets and prevents identically named go_*, process_*, and HTTP metrics from other services from being mixed in.

Prometheus also adds the technical job label automatically. Its value depends on target discovery and is not part of the OSA Proxy contract; queries and rules should not require a specific job value.

Static targets

If one Prometheus server scrapes multiple instances, use one target group. Prometheus assigns a different instance to every address:

scrape_configs:
  - job_name: osa-proxy
    metrics_path: /metrics
    scrape_interval: 30s
    scrape_timeout: 10s
    static_configs:
      - targets:
          - osa-proxy-1.example.com:8080
          - osa-proxy-2.example.com:8080
        labels:
          environment: production
          service: osa-proxy

The resulting series differ by instance:

instance="osa-proxy-1.example.com:8080"
instance="osa-proxy-2.example.com:8080"

Do not expose /metrics to the internet. The endpoint is intended for an internal monitoring system.

HTTP outcome semantics

User traffic and operational probes are separated:

  • /healthz, /readyz, and /metrics are recorded in http_operational_requests_total and excluded from user-traffic SLIs;
  • outcome="success" represents a successful user request;
  • outcome="blocked" represents a policy decision and is not a technical service failure;
  • outcome="client_error" represents an invalid client request;
  • outcome="error" represents an OSA Proxy or dependency failure.

For upstream registries, an expected HTTP 404 has the not_found outcome, other 4xx responses use client_error, and transport errors or 5xx responses use error. Build error ratios from outcome="error", not from all 4xx and 5xx responses combined.

Application metric contract

The tables below list the OSA Proxy-specific metrics. The separately registered promhttp_metric_handler_errors_total and the standard Go/process collector families are described under Runtime and metrics endpoint. For a histogram, Prometheus automatically creates series with the _bucket, _sum, and _count suffixes; they are parts of one metric and are not listed separately. A counter whose name ends in _total is exposed under that same name.

HTTP, CodeScoring, and upstream

MetricTypeLabelsDescription
osa_proxy_infoGauge—Identifies an OSA Proxy target; always equals 1
http_server_requests_secondsHistogramroute, method, status, outcomeUser HTTP request count and latency
http_operational_requests_totalCounterroute, method, status, outcomeCalls to /healthz, /readyz, and /metrics
http_server_active_requestsGauge—Requests currently handled by a replica
gateway_route_requests_secondsHistogrampackage_type, operation, method, repository, status, outcomeFinal upstream registry call count and latency
codescoring_api_requests_secondsHistogramendpoint, method, status, outcome, error_typeCodeScoring API call count and complete latency
codescoring_retry_attempts_totalCounterendpoint, attempt, outcomeCodeScoring API retry attempts
codescoring_circuit_breaker_stateGaugegroup, stateCurrent one-hot circuit-breaker state
codescoring_circuit_breaker_state_changes_totalCountergroup, from, toCircuit-breaker state transitions
codescoring_circuit_breaker_requests_totalCountergroup, outcomeOutcomes of circuit-breaker-protected operations
codescoring_fallbacks_totalCounterendpoint, resultallow or block decisions after CodeScoring failures

Scanner, manifests, and cache

MetricTypeLabelsDescription
scanner_evaluations_totalCounteroperation, outcome, resultBusiness outcomes of package, image, and manifest evaluations
verdict_cache_lookup_purls_totalCounterpackage_type, resultCache hit, miss, and stale results per PURL
verdict_cache_operations_secondsHistogramoperation, package_type, outcomeRedis verdict-cache operation count and latency
manifest_processing_secondsHistogramphase, outcome, resultManifest processing phase duration
manifest_packages_totalCounterresultPackage-level manifest processing results
cache_proactive_refresh_secondsHistogramscope, outcomeProactive refresh cycle and group duration
cache_proactive_refresh_groups_totalCounteroutcomeRefresh-group count by outcome
cache_proactive_refresh_entries_totalCounter—Successfully refreshed cache entries

Artifactory and Nexus discovery

Each provider exports the same metric groups:

MetricsTypeLabelsDescription
artifactory_discovery_syncs_total, nexus_discovery_syncs_totalCounterresultSynchronization-attempt outcomes
artifactory_discovery_sync_duration_seconds, nexus_discovery_sync_duration_secondsHistogramresultSync duration
artifactory_discovery_last_success_timestamp_seconds, nexus_discovery_last_success_timestamp_secondsGauge—Last successful snapshot time
artifactory_discovery_active_repositories, nexus_discovery_active_repositoriesGaugepackage_typeActive dynamic repositories by package type
artifactory_discovery_repository_changes_total, nexus_discovery_repository_changes_totalCounterchangeReconciliation add, update, remove, and skip counts

Managed configuration

MetricTypeLabelsDescription
managed_configuration_desired_revisionGauge—Desired configuration revision
managed_configuration_active_revisionGauge—Revision active on this replica
managed_configuration_convergence_lagGauge—Difference between desired and active revisions
managed_configuration_activation_duration_secondsHistogramstage, result, failureBuild and activation duration
managed_configuration_activation_failures_totalCounterfailureActivation failures by safe category
managed_configuration_store_stateGaugestateOne-hot Configuration Store state
managed_configuration_store_durability_stateGaugestateOne-hot startup durability verification result
managed_configuration_readinessGauge—Whether a valid runtime generation is available
managed_configuration_lease_healthyGauge—Whether the replica lease heartbeat is healthy
managed_configuration_discovery_healthyGaugeproviderDiscovery health for Artifactory and Nexus

Revision gauges show whether a replica has activated the desired state; managed_configuration_convergence_lag should return to 0. Readiness and health gauges use 1 for healthy/available and 0 otherwise. Store state is a one-hot enum with disabled, healthy, degraded, and never_loaded values. Durability state is a one-hot enum with disabled, verified, unverified, and unsafe_accepted values. Activation labels use stage="build|activation", result="success|failure", and failure="none|build|start|activation". Discovery health uses provider="artifactory|nexus".

Lease and managed-discovery health are relevant only when Configuration Store is enabled. The bundled baseline rules do not currently define alerts for these managed-configuration series; add deployment-specific rules when they are required by the operational policy.

Runtime and metrics endpoint

The standard Go collectors provide go_* and process_* metrics for CPU, RSS, heap, GC, goroutines, and file descriptors. promhttp_metric_handler_errors_total{cause} is the standard counter of /metrics gathering or encoding errors.

OSA Proxy does not export network byte counts or network-interface utilization. In Kubernetes, use kubelet/cAdvisor metrics such as container_network_receive_bytes_total and container_network_transmit_bytes_total, filtered by namespace and pod.

When series may be absent

A missing series does not always indicate a failure:

  • whitelist mode does not generate CodeScoring API, retry, circuit-breaker, or fallback traffic;
  • cache metrics appear when Redis verdict caching is enabled and operations occur;
  • proactive refresh metrics require the background task to be enabled;
  • Artifactory and Nexus discovery metrics require the provider to be enabled and at least one sync attempt;
  • histogram series appear after the first matching observation.

Do not replace missing latency with zero because zero latency represents a completed instantaneous request. Return a zero error ratio only when traffic exists but no errors occurred.

PromQL examples

User request rate by replica and outcome:

sum by (environment, instance, outcome) (
  rate(http_server_requests_seconds_count{service="osa-proxy"}[5m])
)

Inbound technical error ratio:

sum by (environment, instance) (
  rate(http_server_requests_seconds_count{service="osa-proxy",outcome="error"}[5m])
)
/
clamp_min(
  sum by (environment, instance) (
    rate(http_server_requests_seconds_count{service="osa-proxy"}[5m])
  ),
  0.001
)

Successful request p95 latency:

histogram_quantile(
  0.95,
  sum by (le, environment, instance) (
    rate(http_server_requests_seconds_bucket{
      service="osa-proxy",
      outcome="success"
    }[5m])
  )
)

Age of the last successful Artifactory discovery:

time() - artifactory_discovery_last_success_timestamp_seconds{service="osa-proxy"}

Recording-rule metrics

When recording rules are loaded, Prometheus calculates seven additional series. These are not application metrics: they are absent from /metrics and exist only after the recording rules have been loaded.

SeriesDimensionsValue
osa_proxy:http_request_rate:5menvironment, instance, outcomeFive-minute user request rate
osa_proxy:http_error_ratio:5menvironment, instanceInbound technical error ratio
osa_proxy:http_latency_p95_seconds:5menvironment, instanceSuccessful inbound request p95 latency
osa_proxy:codescoring_error_ratio:5menvironment, instanceCodeScoring API error ratio
osa_proxy:codescoring_latency_p95_seconds:5menvironment, instanceCodeScoring API p95 latency
osa_proxy:upstream_error_ratio:5menvironment, instanceUpstream registry error ratio
osa_proxy:upstream_latency_p95_seconds:5menvironment, instanceUpstream registry p95 latency

Baseline alerts

Recording rules and baseline alerts can cover the main failure classes:

GroupExample signals
AvailabilityScrape target down, /metrics failures, operational endpoint failures
Inbound trafficHigh technical error ratio, high p95 latency, prolonged no traffic, excessive active requests
CodeScoringAPI errors and latency, retry amplification, open circuit breaker, allow or block fallback
Upstream registriesHigh error ratio/latency, persistent authentication and rate-limit responses
Cache and manifestsRedis cache, scanner, manifest processing, and proactive refresh errors
DiscoverySync failures and stale Artifactory/Nexus last-success snapshots
RuntimeHigh CPU, RSS, GC rate, goroutine count, and file-descriptor usage

Adapt the example thresholds to the actual load. Disable or adjust the no-traffic alert for environments where idle periods are expected.

Pay particular attention to codescoring_fallbacks_total{result="allow"}: even a single fail-open decision may be significant. Detect it with increase(...[5m]) > 0, rather than only a long-running rate() > 0 condition with a large for duration.

Exemplars and tracing

When OpenTelemetry tracing is enabled, histogram metrics may include exemplars with a trace_id. If Prometheus and Grafana are configured to store and display exemplars, operators can navigate from a latency point to the corresponding trace.

Never add trace_id as a regular metric label because that creates a separate series for every request.

Post-installation validation

  1. Check liveness and readiness:

    curl http://localhost:8080/healthz
    curl http://localhost:8080/readyz
  2. Check the metrics endpoint:

    curl http://localhost:8080/metrics
  3. Check the identity metric:

    curl -s http://localhost:8080/metrics | grep '^osa_proxy_info'
  4. Open the Prometheus Targets page and verify that OSA Proxy is UP.

  5. Run the osa_proxy_info query and verify the service, environment, and instance labels. The technical job label is also present, but its specific value is not part of the OSA Proxy configuration contract.

  6. Verify that Prometheus loaded the recording/alerting rules and that Alertmanager accepts a test alert.

Was this page helpful?