10 KiB
Performance Tuning Guide
This document provides guidelines for optimizing the performance of the CAATSM application.
Performance Metrics
Key performance indicators to monitor:
- Message throughput: Messages processed per second
- Latency: End-to-end processing time (NATS receive → DB insert → publish)
- Database query time: Time spent on database operations
- Memory usage: Application memory consumption
- CPU usage: CPU utilization
- Connection pool utilization: Database and NATS connection usage
Batch Processing Configuration
Current Implementation
The application processes messages in batches for efficiency:
[app]
batch_size = 50 # Number of messages per batch
batch_timeout = "2s" # Maximum wait time for a batch
monitor_interval = "30s" # Consumer metrics reporting interval
Tuning Guidelines
Batch Size:
- Small batches (10-50): Lower latency, higher overhead
- Medium batches (50-200): Balanced latency and throughput
- Large batches (200-1000): Higher throughput, higher latency
Recommendations:
- Development: 10-50 messages (faster feedback)
- Production (low latency): 50-100 messages
- Production (high throughput): 200-500 messages
Batch Timeout:
- Low latency: 500ms-1s (process quickly even with small batches)
- Balanced: 2-5s (good balance)
- High throughput: 5-10s (wait for larger batches)
Example Configuration
# High-throughput production configuration
[app]
batch_size = 200
batch_timeout = "5s"
monitor_interval = "30s"
# Low-latency production configuration
[app]
batch_size = 50
batch_timeout = "1s"
monitor_interval = "30s"
Database Connection Pooling
Configuration
[postgres]
url = "postgres://user:pass@db:5432/aviation?sslmode=require"
max_conns = 20 # Maximum connections in pool
min_conns = 5 # Minimum connections in pool
Tuning Guidelines
Connection Pool Size:
- Formula:
max_conns = (expected_concurrent_requests * avg_query_time) / target_latency - Minimum: 2-5 connections (small deployments)
- Recommended: 10-20 connections (medium deployments)
- Maximum: 50-100 connections (high-throughput deployments)
Considerations:
- Each connection consumes memory (~2-5MB)
- PostgreSQL has a maximum connection limit (default: 100)
- Too many connections can degrade performance
- Use connection pooler (PgBouncer) for high concurrency
Example Configurations
# Small deployment (single instance)
[postgres]
max_conns = 10
min_conns = 2
# Medium deployment (2-3 instances)
[postgres]
max_conns = 20
min_conns = 5
# Large deployment (5+ instances, use PgBouncer)
[postgres]
max_conns = 10 # Per instance
min_conns = 2
# Use PgBouncer with pool_mode=transaction
Connection Pool Monitoring
Monitor connection pool metrics:
- Active connections
- Idle connections
- Connection wait time
- Connection errors
Database Indexing Strategy
Current Indexes
The application creates indexes on key fields:
CREATE INDEX idx_telegrams_message_id ON aviation.telegrams (message_id);
CREATE INDEX idx_telegrams_date_time ON aviation.telegrams (date_time);
CREATE INDEX idx_telegrams_priority_indicator ON aviation.telegrams (priority_indicator);
CREATE INDEX idx_telegrams_primary_address ON aviation.telegrams (primary_address);
CREATE INDEX idx_telegrams_received_at ON aviation.telegrams (received_at);
CREATE INDEX idx_telegrams_uuid ON aviation.telegrams (uuid);
Index Optimization
Query Patterns:
- Time-range queries: Index on
received_at(already exists) - Message lookup: Index on
message_id(already exists) - Category filtering: Consider index on
categoryif frequently queried - Composite indexes: For multi-column queries
Example Composite Index:
-- For queries filtering by category and date range
CREATE INDEX idx_telegrams_category_received_at
ON aviation.telegrams (category, received_at DESC);
Index Maintenance
- Monitor index usage: Use
pg_stat_user_indexesto identify unused indexes - Rebuild indexes: Periodically rebuild indexes to reduce bloat
- Concurrent creation: Use
CREATE INDEX CONCURRENTLYin production
NATS JetStream Performance Tuning
Stream Configuration
[nats.stream_limits]
max_msgs = 1000000 # Maximum messages in stream
max_bytes = 1073741824 # Maximum size (1GB)
max_age = "168h" # Retention period (7 days)
discard = "old" # Discard policy
storage = "file" # Storage type (file or memory)
replicas = 3 # Number of replicas
Tuning Guidelines
Storage Type:
- File storage: Persistent, slower (recommended for production)
- Memory storage: Faster, ephemeral (suitable for high-throughput temporary streams)
Replicas:
- Single node: 1 replica (development)
- Production: 3+ replicas (high availability)
Retention:
- Short retention: Lower storage, faster cleanup
- Long retention: More storage, replay capability
Consumer Configuration
[nats.consumer_rules]
max_deliver = 5 # Maximum redelivery attempts
ack_wait = "30s" # ACK wait time
max_ack_pending = 1024 # Maximum unacknowledged messages
deliver_policy = "new" # Delivery policy
backoff = ["5s", "30s", "2m"] # Retry delays
Tuning:
- ack_wait: Set based on processing time (processing_time * 2-3)
- max_ack_pending: Increase for high-throughput (1024-4096)
- backoff: Adjust based on failure patterns
Memory Optimization
Garbage Collection Tuning
Set Go GC environment variables for production:
# Balanced GC (default)
export GOGC=100
# Aggressive GC (lower memory, higher CPU)
export GOGC=50
# Conservative GC (higher memory, lower CPU)
export GOGC=200
Memory Profiling
Use pprof to identify memory issues:
# Enable memory profiling
go tool pprof http://localhost:2112/debug/pprof/heap
# Generate memory profile
go tool pprof -alloc_space http://localhost:2112/debug/pprof/heap
CPU Optimization
Goroutine Management
- Limit goroutines: Use worker pools for concurrent processing
- Context cancellation: Properly cancel goroutines to prevent leaks
- Monitor goroutine count: Use
runtime.NumGoroutine()
CPU Profiling
# Enable CPU profiling
go tool pprof http://localhost:2112/debug/pprof/profile
# 30-second CPU profile
go tool pprof http://localhost:2112/debug/pprof/profile?seconds=30
Monitoring and Profiling
Prometheus Metrics
Key metrics to monitor:
caatsm_messages_total: Message throughputcaatsm_handle_latency_seconds: Processing latencycaatsm_db_query_latency_seconds: Database query timecaatsm_nats_consumer_pending_messages: Consumer lag
Grafana Dashboards
Create dashboards for:
- Message throughput over time
- Latency percentiles (P50, P95, P99)
- Error rates
- Resource utilization (CPU, memory, connections)
Profiling Endpoints
The application exposes profiling endpoints (if enabled):
# Heap profile
curl http://localhost:2112/debug/pprof/heap > heap.prof
# CPU profile
curl http://localhost:2112/debug/pprof/profile?seconds=30 > cpu.prof
# Goroutine profile
curl http://localhost:2112/debug/pprof/goroutine > goroutine.prof
Performance Testing
Load Testing
Use tools like k6, wrk, or vegeta for load testing:
# Example: Generate load with seed-telegrams
go run ./cmd/seed-telegrams \
--count=10000 \
--mode=burst \
--category=mixed
Benchmark Tests
Run built-in benchmarks:
# Run all benchmarks
go test -bench=. -benchmem ./...
# Run specific benchmark
go test -bench=BenchmarkParser -benchmem ./internal/adapter/parser
Performance Baselines
Establish performance baselines:
- Throughput: Messages per second
- Latency: P50, P95, P99 percentiles
- Resource usage: CPU, memory, connections
Optimization Checklist
Application Level
- Optimize batch size for workload
- Tune connection pool sizes
- Review and optimize database queries
- Add missing indexes for query patterns
- Enable query result caching where appropriate
Infrastructure Level
- Use connection pooler (PgBouncer) for high concurrency
- Configure database connection limits appropriately
- Use read replicas for query-heavy workloads
- Optimize NATS JetStream stream configuration
- Scale horizontally (multiple instances)
Monitoring
- Set up performance dashboards
- Configure alerts for performance degradation
- Regular performance profiling
- Monitor resource utilization
- Track performance trends over time
Troubleshooting Performance Issues
High Latency
Symptoms:
- Slow message processing
- High P95/P99 latencies
Investigation:
- Check database query times
- Review NATS consumer lag
- Profile CPU and memory usage
- Check for connection pool exhaustion
Solutions:
- Optimize slow database queries
- Increase batch size
- Add database indexes
- Scale horizontally
Low Throughput
Symptoms:
- Low messages per second
- High CPU usage
Investigation:
- Check for bottlenecks (DB, NATS, CPU)
- Review batch processing configuration
- Profile application code
Solutions:
- Increase batch size
- Optimize hot code paths
- Scale horizontally
- Use connection pooling
High Memory Usage
Symptoms:
- Memory leaks
- High memory consumption
Investigation:
- Heap profiling
- Check for goroutine leaks
- Review batch sizes
Solutions:
- Fix memory leaks
- Reduce batch sizes
- Tune GC settings
- Limit concurrent operations