Skip to content

Monitoring API

Doc Version: 1.0.2
Last Updated: 2026-05-21
Git Commit: 61384b4a
Author: Lincoln

JAiRouter provides comprehensive monitoring APIs for health checks, metrics collection, and system status monitoring.

Overview

The Monitoring API includes:

  • Health Checks - Service and instance health status
  • Metrics - Performance and usage statistics
  • System Status - Overall system health and configuration

All monitoring endpoints are available under the /actuator/* path and provide real-time insights into your JAiRouter deployment.

Health Check Endpoints

System Health

Get overall system health status:

GET /actuator/health

Response:

{
  "status": "UP",
  "components": {
    "diskSpace": {
      "status": "UP",
      "details": {
        "total": 499963174912,
        "free": 91943821312,
        "threshold": 10485760,
        "exists": true
      }
    },
    "ping": {
      "status": "UP"
    },
    "tracing": {
      "status": "UP"
    },
    "sharedState": {
      "status": "UP"
    },
    "rbacEndpointCoverage": {
      "status": "UP",
      "details": {
        "missing": 0
      }
    }
  }
}

⚠️ Component names are the registered health indicators: tracing (TracingHealthIndicator.java:28), sharedState (SharedStateHealthIndicator.java:31), rbacEndpointCoverage (RbacEndpointCoverageHealthIndicator.java:27) and jwtRedisHealthIndicator (JwtRedisHealthMonitor.java:21), plus Spring Boot's standard diskSpace etc. There is no modelRouter component; the activeInstances / totalInstances / circuitBreakerStatus structure does not exist in the code.

Detailed Health Information

Get detailed health information including all components:

GET /actuator/health

⚠️ There is no /actuator/health/detailed endpoint: only the liveness and readiness groups are defined (config/base/monitoring-base.yml:70-84). To see every component, call GET /actuator/health under the dev profile (show-details: always).

Response:

{
  "status": "UP",
  "components": {
    "diskSpace": {
      "status": "UP",
      "details": {
        "total": 499963174912,
        "free": 91943821312,
        "threshold": 10485760,
        "exists": true
      }
    }
  }
}

⚠️ The original document also listed a modelRouter component here (with services.*.instances[] detail) — no such component exists in the code, so it has been removed; see the note at the top of this section for the real health components.

Metrics Endpoints

Application Metrics

Get Prometheus-format metrics:

GET /actuator/prometheus

Response:

# HELP jvm_memory_used_bytes The amount of used memory
# TYPE jvm_memory_used_bytes gauge
jvm_memory_used_bytes{area="heap",id="PS Eden Space",} 2.38026752E8
jvm_memory_used_bytes{area="heap",id="PS Survivor Space",} 1048576.0
jvm_memory_used_bytes{area="heap",id="PS Old Gen",} 4.2991616E7

# HELP jairouter_requests_total Total number of requests
# TYPE jairouter_requests_total counter
jairouter_requests_total{service="chat",method="POST",status="200",} 1247.0
jairouter_requests_total{service="chat",method="POST",status="500",} 23.0
jairouter_requests_total{service="chat",method="POST",status="200",} 1156.0
jairouter_requests_total{service="chat",method="POST",status="500",} 18.0

# HELP jairouter_request_duration_seconds Request duration in seconds
# TYPE jairouter_request_duration_seconds histogram
jairouter_request_duration_seconds_bucket{service="chat",method="POST",le="0.1",} 234.0
jairouter_request_duration_seconds_bucket{service="chat",method="POST",le="0.5",} 892.0
jairouter_request_duration_seconds_bucket{service="chat",method="POST",le="1.0",} 1156.0
jairouter_request_duration_seconds_bucket{service="chat",method="POST",le="+Inf",} 1270.0

# HELP jairouter_circuit_breaker_state Circuit breaker state (0=CLOSED, 1=OPEN, 2=HALF_OPEN)
# TYPE jairouter_circuit_breaker_state gauge
jairouter_circuit_breaker_state{service="chat",method="POST",} 0.0
jairouter_circuit_breaker_state{service="chat",method="POST",} 0.0

# HELP jairouter_rate_limit_remaining Rate limit remaining requests
# TYPE jairouter_rate_limit_remaining gauge
jairouter_rate_limit_remaining{service="chat",client_ip="192.168.1.100",} 45.0
jairouter_rate_limit_remaining{service="chat",client_ip="192.168.1.101",} 38.0

Metrics Summary

Get human-readable metrics summary:

GET /actuator/metrics

Response:

{
  "names": [
    "jvm.memory.used",
    "jvm.memory.max", 
    "jvm.gc.pause",
    "http.server.requests",
    "jairouter_requests_total",
    "jairouter_request_duration_seconds",
    "jairouter_circuit_breaker_state",
    "jairouter_rate_limit_remaining",
    "jairouter_loadbalancer_selections_total"
  ]
}

Specific Metric Details

Get details for a specific metric:

GET /actuator/metrics/jairouter_requests_total

Response:

{
  "name": "jairouter_requests_total",
  "description": "Total number of requests processed by model router",
  "baseUnit": null,
  "measurements": [
    {
      "statistic": "COUNT",
      "value": 2647.0
    }
  ],
  "availableTags": [
    {
      "tag": "service",
      "values": ["chat", "embedding", "rerank", "tts", "stt", "imgGen", "imgEdit"]
    },
    {
      "tag": "method", 
      "values": ["POST", "GET"]
    },
    {
      "tag": "status",
      "values": ["200", "400", "500"]
    }
  ]
}

System Information

Application Info

Get application information:

GET /actuator/info

Response:

{
  "app": {
    "name": "JAiRouter",
    "version": "3.2.2",
    "description": "AI Model Service Router and Load Balancer"
  },
  "build": {
    "version": "3.2.2",
    "artifact": "model-router",
    "name": "model-router",
    "group": "org.unreal",
    "time": "2025-08-19T08:15:30.123Z"
  },
  "git": {
    "branch": "main",
    "commit": {
      "id": "3418d3f6",
      "time": "2025-08-19T08:00:00Z"
    }
  },
  "java": {
    "version": "17.0.8",
    "vendor": "Eclipse Adoptium"
  }
}

⚠️ The build / git sections require the build-info goal in pom.xml (not configured today), so the default /actuator/info does not return them; the output above is illustrative. The current project version is 3.2.2.

Environment Information

Get environment and configuration details:

GET /actuator/env

⚠️ env is not in the default exposure list (config/base/monitoring-base.yml:60 exposes only health / info / metrics / prometheus / jairouter-metrics / error-tracking; see application-prod.yml:40 for prod). Add env to management.endpoints.web.exposure.include first if you need it.

Response:

{
  "activeProfiles": ["default"],
  "propertySources": [
    {
      "name": "server.ports",
      "properties": {
        "local.server.port": {
          "value": 8080
        }
      }
    },
    {
      "name": "applicationConfig: [classpath:/application.yml]",
      "properties": {
        "model-router.load-balancer.default-strategy": {
          "value": "ROUND_ROBIN"
        },
        "model-router.rate-limit.default-algorithm": {
          "value": "TOKEN_BUCKET"
        },
        "model-router.circuit-breaker.failure-threshold": {
          "value": 5
        }
      }
    }
  ]
}

⚠️ The output above is illustrative: every configuration key in this project lives under the jairouter.* prefix — there are no model-router.* keys (no match anywhere under src/main/resources).

Custom Monitoring Endpoints

The gateway exposes the following custom monitoring endpoints:

System Health Status

Get the overall system health status:

GET /api/monitoring/health

Monitoring Configuration

Get the current monitoring configuration:

GET /api/monitoring/config

Circuit Breaker Statistics

Get circuit breaker statistics:

GET /api/monitoring/circuit-breaker/stats

Degradation Status

Get the service degradation status:

GET /api/monitoring/degradation/status

Error Statistics

Get error statistics:

GET /api/monitoring/errors/stats

Cache Statistics

Get cache statistics:

GET /api/monitoring/cache/stats

Note: These are the custom monitoring endpoints JAiRouter (currently 3.2.2) actually provides; each maps to MonitoringController.java:65,226,240,268,309,390,426. See the Management API for more endpoints.

Monitoring Integration

Prometheus Integration

JAiRouter exposes metrics in Prometheus format at /actuator/prometheus. Configure Prometheus to scrape these metrics:

# prometheus.yml
scrape_configs:
  - job_name: 'jairouter'
    static_configs:
      - targets: ['localhost:8080']
    metrics_path: '/actuator/prometheus'
    scrape_interval: 15s

Grafana Dashboard

Import the JAiRouter Grafana dashboard for visualization:

  • Dashboard ID: Coming soon
  • Metrics: Request rate, response time, error rate, circuit breaker status
  • Alerts: High error rate, circuit breaker open, instance down

Health Check Integration

Configure external monitoring tools to check JAiRouter health:

# Simple health check
curl -f http://localhost:8080/actuator/health || exit 1

# Detailed health check with specific component
curl -f http://localhost:8080/actuator/health/tracing || exit 1

Monitoring Best Practices

Key Metrics to Monitor

  1. Request Metrics
  2. Request rate (requests/second)
  3. Response time (p50, p95, p99)
  4. Error rate (percentage)

  5. Instance Health

  6. Instance availability
  7. Health check response time
  8. Circuit breaker state

  9. Resource Usage

  10. JVM memory usage
  11. CPU utilization
  12. Disk space

  13. Rate Limiting

  14. Rate limit utilization
  15. Rejected requests
  16. Client-specific limits

Alerting Rules

Set up alerts for critical conditions:

# Prometheus alerting rules
groups:
  - name: jairouter
    rules:
      - alert: HighErrorRate
        expr: rate(jairouter_requests_total{status="500"}[5m]) / rate(jairouter_requests_total[5m]) > 0.1
        for: 2m
        labels:
          severity: warning
        annotations:
          summary: "High error rate detected"

      - alert: CircuitBreakerOpen
        expr: jairouter_circuit_breaker_state > 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "Circuit breaker is open"

      - alert: InstanceDown
        expr: up{job="jairouter"} == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "JAiRouter instance is down"

Log Monitoring

Monitor application logs for:

  • Error patterns
  • Performance issues
  • Configuration changes
  • Security events
# Tail logs with filtering
tail -f logs/application.log | grep -E "(ERROR|WARN|Circuit|Rate)"

Troubleshooting

Common Issues

  1. High Response Time
  2. Check instance health
  3. Review load balancer distribution
  4. Monitor resource usage

  5. Circuit Breaker Open

  6. Check instance connectivity
  7. Review error logs
  8. Verify instance configuration

  9. Rate Limit Exceeded

  10. Review rate limit configuration
  11. Check client request patterns
  12. Consider increasing limits

Debug Endpoints

Enable debug logging for detailed monitoring:

# application.yml
logging:
  level:
    org.unreal.modelrouter: DEBUG
    org.springframework.web: DEBUG

Access debug information:

GET /actuator/loggers/org.unreal.modelrouter

⚠️ loggers is not in the default exposure list — add it to management.endpoints.web.exposure.include first.

Security Considerations

Monitoring Endpoint Security

Secure monitoring endpoints in production:

# application.yml
management:
  endpoints:
    web:
      exposure:
        include: health,info,prometheus
  endpoint:
    health:
      show-details: when-authorized
  security:
    enabled: true

Sensitive Information

Avoid exposing sensitive data in metrics:

  • API keys
  • Internal URLs
  • User information
  • Configuration secrets

Next Steps