Appearance
Monitoring & Scaling API
Overview
The Monitoring & Scaling API provides real-time insights into system health, performance metrics, and scaling recommendations. This API is essential for operational monitoring, capacity planning, and automated scaling decisions.
Endpoints
GET /api/monitoring/scaling # Get current system metrics
POST /api/monitoring/scaling # Update system configurationAuthentication
Requires valid session authentication with admin privileges. Only SUPER_ADMIN and ACCOUNT_OWNER roles can access monitoring data.
Rate Limits
- All Tiers: 1,000 requests/hour (monitoring endpoints are not rate-limited for operational needs)
Get System Metrics
Endpoint
GET /api/monitoring/scalingResponse Format
json
{
"success": true,
"data": {
"health": "green" | "yellow" | "orange" | "red",
"timestamp": "string (ISO 8601)",
"scaling": {
"metrics": {
"timestamp": "string",
"cpu": {
"usage": number,
"trend": "stable" | "increasing" | "decreasing"
},
"memory": {
"usage": number,
"available": number,
"trend": "stable" | "increasing" | "decreasing"
},
"queue": {
"depth": number,
"processingRate": number,
"avgWaitTime": number,
"byTier": {
"enterprise": {
"depth": number,
"avgWaitTime": number
},
"professional": {
"depth": number,
"avgWaitTime": number
},
"standard": {
"depth": number,
"avgWaitTime": number
},
"free": {
"depth": number,
"avgWaitTime": number
}
}
},
"performance": {
"avgResponseTime": number,
"p95ResponseTime": number,
"errorRate": number,
"throughput": number
},
"cost": {
"currentHourly": number,
"projectedMonthly": number,
"costPerRequest": number
},
"recommendations": {
"action": "scale_up" | "scale_down" | "maintain",
"reason": "string",
"estimatedCostImpact": number,
"estimatedPerformanceGain": number
}
},
"readiness": {
"currentMode": "startup" | "growth" | "enterprise",
"isReady": boolean,
"blockers": ["string"],
"recommendations": ["string"],
"estimatedMonthlyCost": {
"current": number,
"withAWS": number,
"withGCP": number,
"withMultiCloud": number
}
}
},
"queue": {
"total": number,
"byTier": {
"enterprise": {
"queued": number,
"processing": number,
"completed": number,
"failed": number,
"avgWaitTime": number,
"avgProcessingTime": number
},
"professional": {
"queued": number,
"processing": number,
"completed": number,
"failed": number,
"avgWaitTime": number,
"avgProcessingTime": number
},
"standard": {
"queued": number,
"processing": number,
"completed": number,
"failed": number,
"avgWaitTime": number,
"avgProcessingTime": number
},
"free": {
"queued": number,
"processing": number,
"completed": number,
"failed": number,
"avgWaitTime": number,
"avgProcessingTime": number
}
},
"throughput": {
"requestsPerSecond": number,
"requestsPerMinute": number
}
},
"throttling": {
"systemHealth": "green" | "yellow" | "orange" | "red",
"globalTokens": number,
"activeUsers": number,
"providers": [
{
"name": "string",
"circuitState": "closed" | "open" | "half_open",
"failures": number
}
]
},
"alerts": {
"cpu": boolean,
"memory": boolean,
"queue": boolean,
"cost": boolean
}
}
}Update System Configuration
Endpoint
POST /api/monitoring/scalingRequest Body
json
{
"action": "updateHealth" | "toggleScaling" | "updateLimits",
"config": {
"health": "green" | "yellow" | "orange" | "red",
"scalingEnabled": boolean,
"costLimits": {
"hourlyBudget": number,
"monthlyBudget": number,
"alertThreshold": number
}
}
}Response
json
{
"success": true,
"message": "Action updateHealth completed successfully"
}Health Status Levels
| Level | Description | CPU Threshold | Memory Threshold | Queue Depth |
|---|---|---|---|---|
| Green | Normal operation | <60% | <60% | <500 |
| Yellow | Moderate load, some delays | 60-80% | 60-80% | 500-1000 |
| Orange | High load, free tier restricted | 80-95% | 80-95% | 1000-2000 |
| Red | Critical load, enterprise only | >95% | >95% | >2000 |
Scaling Recommendations
Action Types
| Action | Trigger | Description |
|---|---|---|
scale_up | High resource usage | Add more capacity |
scale_down | Low resource usage | Reduce capacity to save costs |
maintain | Normal operation | Keep current configuration |
Recommendation Factors
- CPU Usage: Current and trending CPU utilization
- Memory Usage: Current and trending memory utilization
- Queue Depth: Number of pending requests
- Response Time: Average and P95 response times
- Cost Efficiency: Cost per request and budget constraints
- Error Rate: System error rate and reliability
Queue Management
Tier Priority
- Enterprise (Priority 1): 30% reserved capacity, <100ms SLA
- Professional (Priority 2): 40% shared capacity, <500ms SLA
- Standard (Priority 3): 25% shared capacity, <2s SLA
- Free (Priority 4): 5% overflow capacity, best effort
Throttling Behavior
| Health Level | Enterprise | Professional | Standard | Free |
|---|---|---|---|---|
| Green | ✅ Normal | ✅ Normal | ✅ Normal | ✅ Normal |
| Yellow | ✅ Normal | ✅ Normal | ✅ Normal | ⚠️ Slower |
| Orange | ✅ Normal | ✅ Normal | ⚠️ Slower | ❌ Blocked |
| Red | ✅ Normal | ❌ Blocked | ❌ Blocked | ❌ Blocked |
Error Responses
401 Unauthorized
json
{
"error": "Unauthorized",
"message": "Admin privileges required"
}403 Forbidden
json
{
"error": "Forbidden",
"message": "SUPER_ADMIN or ACCOUNT_OWNER role required"
}500 Internal Server Error
json
{
"error": "Failed to fetch monitoring data",
"requestId": "string"
}Example Usage
Get Current Metrics (cURL)
bash
curl -X GET https://app.veriprompt.tech/api/monitoring/scaling \
-H "Authorization: Bearer <admin-token>"JavaScript Dashboard Integration
javascript
// Real-time monitoring dashboard
class MonitoringDashboard {
constructor(token) {
this.token = token;
this.updateInterval = null;
}
async fetchMetrics() {
const response = await fetch('/api/monitoring/scaling', {
headers: {
'Authorization': `Bearer ${this.token}`
}
});
if (!response.ok) {
throw new Error(`HTTP ${response.status}: ${response.statusText}`);
}
return await response.json();
}
startRealTimeUpdates(callback, intervalMs = 5000) {
this.updateInterval = setInterval(async () => {
try {
const metrics = await this.fetchMetrics();
callback(metrics);
} catch (error) {
console.error('Failed to fetch metrics:', error);
}
}, intervalMs);
}
stopRealTimeUpdates() {
if (this.updateInterval) {
clearInterval(this.updateInterval);
this.updateInterval = null;
}
}
}
// Usage
const dashboard = new MonitoringDashboard(adminToken);
dashboard.startRealTimeUpdates((metrics) => {
updateHealthIndicator(metrics.data.health);
updateQueueStats(metrics.data.queue);
updateScalingRecommendations(metrics.data.scaling.metrics.recommendations);
});Python Monitoring Script
python
import requests
import time
import json
from datetime import datetime
class SystemMonitor:
def __init__(self, base_url, token):
self.base_url = base_url
self.headers = {'Authorization': f'Bearer {token}'}
def get_metrics(self):
"""Fetch current system metrics"""
response = requests.get(
f'{self.base_url}/api/monitoring/scaling',
headers=self.headers
)
response.raise_for_status()
return response.json()
def check_alerts(self, metrics):
"""Check for alert conditions"""
alerts = []
data = metrics['data']
if data['alerts']['cpu']:
alerts.append('CPU usage critical')
if data['alerts']['memory']:
alerts.append('Memory usage critical')
if data['alerts']['queue']:
alerts.append('Queue depth critical')
if data['alerts']['cost']:
alerts.append('Cost threshold exceeded')
return alerts
def monitor_continuous(self, interval=60):
"""Continuous monitoring with alerting"""
print("Starting continuous monitoring...")
while True:
try:
metrics = self.get_metrics()
timestamp = datetime.now().strftime('%Y-%m-%d %H:%M:%S')
print(f"\n[{timestamp}] System Status: {metrics['data']['health'].upper()}")
# Check for alerts
alerts = self.check_alerts(metrics)
if alerts:
print("🚨 ALERTS:")
for alert in alerts:
print(f" - {alert}")
# Show key metrics
scaling = metrics['data']['scaling']['metrics']
print(f"CPU: {scaling['cpu']['usage']:.1f}%")
print(f"Memory: {scaling['memory']['usage']:.1f}%")
print(f"Queue: {metrics['data']['queue']['total']} requests")
print(f"Throughput: {metrics['data']['queue']['throughput']['requestsPerSecond']:.1f} req/s")
# Show recommendations
rec = scaling['recommendations']
if rec['action'] != 'maintain':
print(f"💡 Recommendation: {rec['action']} - {rec['reason']}")
time.sleep(interval)
except KeyboardInterrupt:
print("\nMonitoring stopped")
break
except Exception as e:
print(f"Error: {e}")
time.sleep(interval)
# Usage
monitor = SystemMonitor('https://app.veriprompt.tech', admin_token)
monitor.monitor_continuous(interval=30) # Check every 30 secondsCloud Readiness Assessment
Readiness Indicators
The API provides insights into multi-cloud scaling readiness:
json
{
"readiness": {
"currentMode": "startup",
"isReady": false,
"blockers": [
"AWS credentials not configured",
"GCP project not configured"
],
"recommendations": [
"Fix error rate issues before scaling",
"Consider implementing caching before scaling"
],
"estimatedMonthlyCost": {
"current": 100,
"withAWS": 500,
"withGCP": 600,
"withMultiCloud": 1000
}
}
}Preparation Checklist
- [ ] Cloud Credentials: Configure API keys for target providers
- [ ] Network Configuration: Set up VPC and security groups
- [ ] Monitoring Setup: Install monitoring agents
- [ ] Load Balancer: Configure load balancing rules
- [ ] Database Migration: Prepare for distributed database
- [ ] Cost Controls: Set up billing alerts and limits
Best Practices
Monitoring
- Set up alerts: Configure alerts for critical thresholds
- Monitor trends: Track metrics over time, not just point-in-time
- Dashboard automation: Build automated dashboards for key stakeholders
- Historical analysis: Keep historical data for capacity planning
Scaling
- Gradual scaling: Scale up/down gradually to avoid instability
- Test scaling: Test scaling procedures in non-production environments
- Monitor costs: Track actual vs predicted scaling costs
- Rollback plans: Have rollback procedures for failed scaling attempts
Performance
- Baseline metrics: Establish baseline performance metrics
- Capacity planning: Plan capacity based on growth projections
- Load testing: Regular load testing to validate scaling thresholds
- Optimization: Continuously optimize based on monitoring insights
Webhook Integration
For real-time alerting, configure webhooks to receive instant notifications:
json
{
"webhookUrl": "https://your-app.com/alerts",
"events": ["health_degraded", "scaling_recommended", "cost_threshold"],
"headers": {
"Authorization": "Bearer your-webhook-token"
}
}Changelog
Version 1.0.0 (Current)
- Initial release
- Real-time metrics collection
- Multi-tier queue management
- Scaling recommendations
- Cloud readiness assessment
- Alert management
