Skip to content

Monitoring & Scaling API ​

Overview ​

The Monitoring & Scaling API provides real-time insights into system health, performance metrics, and scaling recommendations. This API is essential for operational monitoring, capacity planning, and automated scaling decisions.

Endpoints ​

GET /api/monitoring/scaling     # Get current system metrics
POST /api/monitoring/scaling    # Update system configuration

Authentication ​

Requires valid session authentication with admin privileges. Only SUPER_ADMIN and ACCOUNT_OWNER roles can access monitoring data.

Rate Limits ​

  • All Tiers: 1,000 requests/hour (monitoring endpoints are not rate-limited for operational needs)

Get System Metrics ​

Endpoint ​

GET /api/monitoring/scaling

Response Format ​

json
{
  "success": true,
  "data": {
    "health": "green" | "yellow" | "orange" | "red",
    "timestamp": "string (ISO 8601)",
    "scaling": {
      "metrics": {
        "timestamp": "string",
        "cpu": {
          "usage": number,
          "trend": "stable" | "increasing" | "decreasing"
        },
        "memory": {
          "usage": number,
          "available": number,
          "trend": "stable" | "increasing" | "decreasing"
        },
        "queue": {
          "depth": number,
          "processingRate": number,
          "avgWaitTime": number,
          "byTier": {
            "enterprise": {
              "depth": number,
              "avgWaitTime": number
            },
            "professional": {
              "depth": number,
              "avgWaitTime": number  
            },
            "standard": {
              "depth": number,
              "avgWaitTime": number
            },
            "free": {
              "depth": number,
              "avgWaitTime": number
            }
          }
        },
        "performance": {
          "avgResponseTime": number,
          "p95ResponseTime": number,
          "errorRate": number,
          "throughput": number
        },
        "cost": {
          "currentHourly": number,
          "projectedMonthly": number,
          "costPerRequest": number
        },
        "recommendations": {
          "action": "scale_up" | "scale_down" | "maintain",
          "reason": "string",
          "estimatedCostImpact": number,
          "estimatedPerformanceGain": number
        }
      },
      "readiness": {
        "currentMode": "startup" | "growth" | "enterprise",
        "isReady": boolean,
        "blockers": ["string"],
        "recommendations": ["string"],
        "estimatedMonthlyCost": {
          "current": number,
          "withAWS": number,
          "withGCP": number,
          "withMultiCloud": number
        }
      }
    },
    "queue": {
      "total": number,
      "byTier": {
        "enterprise": {
          "queued": number,
          "processing": number,
          "completed": number,
          "failed": number,
          "avgWaitTime": number,
          "avgProcessingTime": number
        },
        "professional": {
          "queued": number,
          "processing": number,
          "completed": number,
          "failed": number,
          "avgWaitTime": number,
          "avgProcessingTime": number
        },
        "standard": {
          "queued": number,
          "processing": number,
          "completed": number,
          "failed": number,
          "avgWaitTime": number,
          "avgProcessingTime": number
        },
        "free": {
          "queued": number,
          "processing": number,
          "completed": number,
          "failed": number,
          "avgWaitTime": number,
          "avgProcessingTime": number
        }
      },
      "throughput": {
        "requestsPerSecond": number,
        "requestsPerMinute": number
      }
    },
    "throttling": {
      "systemHealth": "green" | "yellow" | "orange" | "red",
      "globalTokens": number,
      "activeUsers": number,
      "providers": [
        {
          "name": "string",
          "circuitState": "closed" | "open" | "half_open",
          "failures": number
        }
      ]
    },
    "alerts": {
      "cpu": boolean,
      "memory": boolean,
      "queue": boolean,
      "cost": boolean
    }
  }
}

Update System Configuration ​

Endpoint ​

POST /api/monitoring/scaling

Request Body ​

json
{
  "action": "updateHealth" | "toggleScaling" | "updateLimits",
  "config": {
    "health": "green" | "yellow" | "orange" | "red",
    "scalingEnabled": boolean,
    "costLimits": {
      "hourlyBudget": number,
      "monthlyBudget": number,
      "alertThreshold": number
    }
  }
}

Response ​

json
{
  "success": true,
  "message": "Action updateHealth completed successfully"
}

Health Status Levels ​

LevelDescriptionCPU ThresholdMemory ThresholdQueue Depth
GreenNormal operation<60%<60%<500
YellowModerate load, some delays60-80%60-80%500-1000
OrangeHigh load, free tier restricted80-95%80-95%1000-2000
RedCritical load, enterprise only>95%>95%>2000

Scaling Recommendations ​

Action Types ​

ActionTriggerDescription
scale_upHigh resource usageAdd more capacity
scale_downLow resource usageReduce capacity to save costs
maintainNormal operationKeep current configuration

Recommendation Factors ​

  • CPU Usage: Current and trending CPU utilization
  • Memory Usage: Current and trending memory utilization
  • Queue Depth: Number of pending requests
  • Response Time: Average and P95 response times
  • Cost Efficiency: Cost per request and budget constraints
  • Error Rate: System error rate and reliability

Queue Management ​

Tier Priority ​

  1. Enterprise (Priority 1): 30% reserved capacity, <100ms SLA
  2. Professional (Priority 2): 40% shared capacity, <500ms SLA
  3. Standard (Priority 3): 25% shared capacity, <2s SLA
  4. Free (Priority 4): 5% overflow capacity, best effort

Throttling Behavior ​

Health LevelEnterpriseProfessionalStandardFree
Green✅ Normal✅ Normal✅ Normal✅ Normal
Yellow✅ Normal✅ Normal✅ Normal⚠️ Slower
Orange✅ Normal✅ Normal⚠️ Slower❌ Blocked
Red✅ Normal❌ Blocked❌ Blocked❌ Blocked

Error Responses ​

401 Unauthorized ​

json
{
  "error": "Unauthorized",
  "message": "Admin privileges required"
}

403 Forbidden ​

json
{
  "error": "Forbidden", 
  "message": "SUPER_ADMIN or ACCOUNT_OWNER role required"
}

500 Internal Server Error ​

json
{
  "error": "Failed to fetch monitoring data",
  "requestId": "string"
}

Example Usage ​

Get Current Metrics (cURL) ​

bash
curl -X GET https://app.veriprompt.tech/api/monitoring/scaling \
  -H "Authorization: Bearer <admin-token>"

JavaScript Dashboard Integration ​

javascript
// Real-time monitoring dashboard
class MonitoringDashboard {
  constructor(token) {
    this.token = token;
    this.updateInterval = null;
  }
  
  async fetchMetrics() {
    const response = await fetch('/api/monitoring/scaling', {
      headers: {
        'Authorization': `Bearer ${this.token}`
      }
    });
    
    if (!response.ok) {
      throw new Error(`HTTP ${response.status}: ${response.statusText}`);
    }
    
    return await response.json();
  }
  
  startRealTimeUpdates(callback, intervalMs = 5000) {
    this.updateInterval = setInterval(async () => {
      try {
        const metrics = await this.fetchMetrics();
        callback(metrics);
      } catch (error) {
        console.error('Failed to fetch metrics:', error);
      }
    }, intervalMs);
  }
  
  stopRealTimeUpdates() {
    if (this.updateInterval) {
      clearInterval(this.updateInterval);
      this.updateInterval = null;
    }
  }
}

// Usage
const dashboard = new MonitoringDashboard(adminToken);
dashboard.startRealTimeUpdates((metrics) => {
  updateHealthIndicator(metrics.data.health);
  updateQueueStats(metrics.data.queue);
  updateScalingRecommendations(metrics.data.scaling.metrics.recommendations);
});

Python Monitoring Script ​

python
import requests
import time
import json
from datetime import datetime

class SystemMonitor:
    def __init__(self, base_url, token):
        self.base_url = base_url
        self.headers = {'Authorization': f'Bearer {token}'}
        
    def get_metrics(self):
        """Fetch current system metrics"""
        response = requests.get(
            f'{self.base_url}/api/monitoring/scaling',
            headers=self.headers
        )
        response.raise_for_status()
        return response.json()
    
    def check_alerts(self, metrics):
        """Check for alert conditions"""
        alerts = []
        data = metrics['data']
        
        if data['alerts']['cpu']:
            alerts.append('CPU usage critical')
        if data['alerts']['memory']: 
            alerts.append('Memory usage critical')
        if data['alerts']['queue']:
            alerts.append('Queue depth critical')
        if data['alerts']['cost']:
            alerts.append('Cost threshold exceeded')
            
        return alerts
    
    def monitor_continuous(self, interval=60):
        """Continuous monitoring with alerting"""
        print("Starting continuous monitoring...")
        
        while True:
            try:
                metrics = self.get_metrics()
                timestamp = datetime.now().strftime('%Y-%m-%d %H:%M:%S')
                
                print(f"\n[{timestamp}] System Status: {metrics['data']['health'].upper()}")
                
                # Check for alerts
                alerts = self.check_alerts(metrics)
                if alerts:
                    print("🚨 ALERTS:")
                    for alert in alerts:
                        print(f"  - {alert}")
                
                # Show key metrics
                scaling = metrics['data']['scaling']['metrics']
                print(f"CPU: {scaling['cpu']['usage']:.1f}%")
                print(f"Memory: {scaling['memory']['usage']:.1f}%")  
                print(f"Queue: {metrics['data']['queue']['total']} requests")
                print(f"Throughput: {metrics['data']['queue']['throughput']['requestsPerSecond']:.1f} req/s")
                
                # Show recommendations
                rec = scaling['recommendations']
                if rec['action'] != 'maintain':
                    print(f"💡 Recommendation: {rec['action']} - {rec['reason']}")
                
                time.sleep(interval)
                
            except KeyboardInterrupt:
                print("\nMonitoring stopped")
                break
            except Exception as e:
                print(f"Error: {e}")
                time.sleep(interval)

# Usage
monitor = SystemMonitor('https://app.veriprompt.tech', admin_token)
monitor.monitor_continuous(interval=30)  # Check every 30 seconds

Cloud Readiness Assessment ​

Readiness Indicators ​

The API provides insights into multi-cloud scaling readiness:

json
{
  "readiness": {
    "currentMode": "startup",
    "isReady": false,
    "blockers": [
      "AWS credentials not configured",
      "GCP project not configured"
    ],
    "recommendations": [
      "Fix error rate issues before scaling",
      "Consider implementing caching before scaling"
    ],
    "estimatedMonthlyCost": {
      "current": 100,
      "withAWS": 500,
      "withGCP": 600, 
      "withMultiCloud": 1000
    }
  }
}

Preparation Checklist ​

  • [ ] Cloud Credentials: Configure API keys for target providers
  • [ ] Network Configuration: Set up VPC and security groups
  • [ ] Monitoring Setup: Install monitoring agents
  • [ ] Load Balancer: Configure load balancing rules
  • [ ] Database Migration: Prepare for distributed database
  • [ ] Cost Controls: Set up billing alerts and limits

Best Practices ​

Monitoring ​

  1. Set up alerts: Configure alerts for critical thresholds
  2. Monitor trends: Track metrics over time, not just point-in-time
  3. Dashboard automation: Build automated dashboards for key stakeholders
  4. Historical analysis: Keep historical data for capacity planning

Scaling ​

  1. Gradual scaling: Scale up/down gradually to avoid instability
  2. Test scaling: Test scaling procedures in non-production environments
  3. Monitor costs: Track actual vs predicted scaling costs
  4. Rollback plans: Have rollback procedures for failed scaling attempts

Performance ​

  1. Baseline metrics: Establish baseline performance metrics
  2. Capacity planning: Plan capacity based on growth projections
  3. Load testing: Regular load testing to validate scaling thresholds
  4. Optimization: Continuously optimize based on monitoring insights

Webhook Integration ​

For real-time alerting, configure webhooks to receive instant notifications:

json
{
  "webhookUrl": "https://your-app.com/alerts",
  "events": ["health_degraded", "scaling_recommended", "cost_threshold"],
  "headers": {
    "Authorization": "Bearer your-webhook-token"
  }
}

Changelog ​

Version 1.0.0 (Current) ​

  • Initial release
  • Real-time metrics collection
  • Multi-tier queue management
  • Scaling recommendations
  • Cloud readiness assessment
  • Alert management