11 KiB
Production Deployment Guide
This guide covers deploying and running the CAATSM application in production environments.
Overview
Production deployments require:
- JetStream mode (mandatory) - for message reliability and persistence
- Manual Stream/Consumer creation - not auto-created in production
- Proper configuration - using
configs/config.prod.toml - SSL/TLS connections - for secure database and telemetry connections
- High availability - stream replicas set to 3+ for HA
Prerequisites
Before deploying to production:
-
Build the binary:
make build # Binary will be at bin/receiver -
NATS JetStream cluster - must be running and accessible
-
PostgreSQL/TimescaleDB - database must be accessible with SSL
-
OpenTelemetry collector (optional) - for observability
-
Prometheus (optional) - for metrics scraping
Configuration
1. Production Configuration File
The repository includes a production configuration template at configs/config.prod.toml. Copy and customize it:
cp configs/config.prod.toml configs/config.prod.toml.local
# Edit config.prod.toml.local with your production values
Key configuration sections:
[nats]
url = "nats://nats.prod:4222"
mode = "jetstream" # Production MUST use JetStream
stream = "TELEGRAM"
consumer = "telegram-consumer"
[nats.stream_limits]
max_msgs = 1000000 # Adjust based on requirements
max_bytes = 1073741824 # 1GB
max_age = "168h" # 7 days
discard = "old"
storage = "file"
replicas = 3 # Use 3+ for HA in production
[nats.consumer_rules]
max_deliver = 5
ack_wait = "30s"
max_ack_pending = 1024
deliver_policy = "new" # Start from new messages in production
replay_policy = "instant"
backoff = ["5s", "30s", "2m", "5m"]
[subscription]
topic = "telegram.serial"
queue_group = "tele-queue"
[postgres]
url = "postgres://user:password@db.prod:5432/aviation?sslmode=require"
max_conns = 20
min_conns = 5
[app]
batch_size = 100
batch_timeout = "2s"
monitor_interval = "30s"
[log]
level = "info"
format = "json"
output = ["stdout"]
[telemetry]
enabled = true
endpoint = "otel-collector.prod:4318"
insecure = false # Use TLS in production
[monitoring]
disabled = false
addr = ":2112"
enable_metrics = true
enable_health = true
[dlq]
enabled = true
subject = "caatsm.dlq"
2. Environment Variables
Alternatively, you can use environment variables instead of a config file:
export GO_ENV=prod
export CAATSM_NATS_URL=nats://nats.prod:4222
export CAATSM_NATS_MODE=jetstream
export CAATSM_POSTGRES_URL=postgres://user:pass@db:5432/aviation?sslmode=require
export CAATSM_LOG_LEVEL=info
export CAATSM_TELEMETRY_ENABLED=true
export CAATSM_TELEMETRY_ENDPOINT=otel-collector.prod:4318
export CAATSM_TELEMETRY_INSECURE=false
JetStream Setup
Pre-create Stream and Consumer
IMPORTANT: The application does NOT auto-create streams/consumers in production. You must create them manually before starting the application.
Using NATS CLI
# Create stream
nats stream add TELEGRAM \
--subjects "telegram.serial,telegram.json" \
--storage file \
--replicas 3 \
--max-msgs 1000000 \
--max-bytes 1GB \
--max-age 7d \
--discard old
# Create consumer
nats consumer add TELEGRAM telegram-consumer \
--filter "telegram.serial" \
--ack explicit \
--deliver new \
--max-deliver 5 \
--ack-wait 30s \
--max-pending 1024
Using NATS Management API
You can also create streams/consumers programmatically using the NATS management API or configuration files.
Verify Setup
# Check stream exists
nats stream info TELEGRAM
# Check consumer exists
nats consumer info TELEGRAM telegram-consumer
# Test connection
nats pub telegram.serial "ZCZC TEST 150631..."
Running the Application
Using Make
make run-prod # GO_ENV=prod (requires config.prod.toml)
Using Task
task run-prod # GO_ENV=prod (requires config.prod.toml)
Direct Execution
# Using binary with config file
GO_ENV=prod ./bin/receiver listen
# Or with environment variables (no config file needed)
GO_ENV=prod \
CAATSM_NATS_URL=nats://nats.prod:4222 \
CAATSM_NATS_MODE=jetstream \
CAATSM_POSTGRES_URL=postgres://user:pass@db:5432/aviation?sslmode=require \
./bin/receiver listen
Production Checklist
Before deploying to production, verify:
- ✅
nats.mode = "jetstream"in config (mandatory) - ✅ Stream and Consumer created manually
- ✅ Stream replicas set to 3+ for high availability
- ✅ PostgreSQL connection configured with SSL (
sslmode=require) - ✅ Log level set to
infoorwarn(notdebug) - ✅ Log format set to
jsonfor log aggregation - ✅ Telemetry endpoint configured (if using observability)
- ✅ Telemetry TLS enabled (
insecure = false) - ✅ Monitoring endpoints exposed for Prometheus scraping
- ✅ DLQ enabled for poison message handling
- ✅ Appropriate retention limits configured (max_msgs, max_bytes, max_age)
- ✅ Connection pool sizes appropriate for workload
- ✅ Batch sizes tuned for throughput
Monitoring and Observability
Health Endpoints
The application exposes health and readiness endpoints:
GET /livez- Liveness endpoint (process health)GET /readyz- Readiness endpoint (checks PostgreSQL and NATS)GET /healthz- Alias for/readyzGET /metrics- Prometheus metrics
Configure Prometheus to scrape metrics:
scrape_configs:
- job_name: 'caatsm'
static_configs:
- targets: ['caatsm:2112']
Key Metrics
Monitor these metrics in production:
caatsm_messages_total{stream,consumer,result}- Message throughput and resultscaatsm_handle_latency_seconds_bucket- Processing latencycaatsm_retries_total{stream,consumer,reason}- Retry countscaatsm_db_queries_total{operation,result}- Database activitycaatsm_dlq_messages_total- Dead-letter queue messagescaatsm_nats_consumer_pending_messages- Consumer backlog/lag
Logging
Production logs are in JSON format for easy parsing by log aggregation systems:
{
"level": "info",
"ts": 1234567890.123,
"caller": "nats/consumer.go:123",
"msg": "Started consuming messages",
"subject": "telegram.serial",
"consumer": "telegram-consumer",
"stream": "TELEGRAM"
}
High Availability
Multiple Instances
Run multiple instances of the application for high availability:
- All instances use the same durable consumer name
- JetStream distributes messages across instances
- Each instance independently fetches messages
- If an instance fails, others continue processing
Stream Replication
Configure stream with 3+ replicas for HA:
[nats.stream_limits]
replicas = 3 # Minimum 3 for HA, 5 for better distribution
Database Connection Pooling
Configure appropriate connection pool sizes:
[postgres]
max_conns = 20 # Adjust based on number of instances
min_conns = 5
Troubleshooting
Stream Not Found
Error: stream TELEGRAM not found
Solution: Create the stream manually before starting the application (see "JetStream Setup" above).
Consumer Not Found
Error: consumer telegram-consumer not found in stream TELEGRAM
Solution: Create the consumer manually before starting the application (see "JetStream Setup" above).
Messages Not Being Consumed
Symptoms: High pending count, no messages processed
Check:
- Verify consumer exists:
nats consumer info TELEGRAM telegram-consumer - Check pending messages:
nats consumer next TELEGRAM telegram-consumer - Verify application is running and connected
- Check logs for errors
Solutions:
- Increase
batch_sizeif processing is slow - Add more consumer instances
- Check for processing errors in logs
High Pending Count
Symptoms: Consumer has many pending messages
Solutions:
- Increase
batch_sizein config - Add more application instances
- Check processing latency
- Verify database performance
Messages Being Redelivered
Symptoms: Same messages processed multiple times
Check:
- Processing logs for errors
ack_waittimeout may be too short- Processing may be taking longer than
ack_wait
Solutions:
- Increase
ack_waitif processing takes longer - Fix processing errors
- Check database connection and performance
Connection Issues
NATS Connection:
- Verify NATS server is accessible
- Check network connectivity
- Verify NATS URL in config
PostgreSQL Connection:
- Verify database is accessible
- Check SSL certificate configuration
- Verify connection string format
- Check firewall rules
Deployment Options
Systemd Deployment
See docs/deploy-systemd.md for a complete systemd service deployment example.
Kubernetes Deployment
See docs/deploy-k8s.md for Kubernetes deployment with ConfigMap/Secret and health probes.
Performance Tuning
Batch Processing
Adjust batch size based on message size and processing time:
[app]
batch_size = 100 # Increase for higher throughput
batch_timeout = "2s" # Adjust based on latency requirements
Connection Pools
Tune database connection pool:
[postgres]
max_conns = 20 # Total connections across all instances
min_conns = 5 # Keep-alive connections
Stream Retention
Configure retention based on requirements:
[nats.stream_limits]
max_msgs = 1000000 # Maximum messages
max_bytes = 1073741824 # Maximum size (1GB)
max_age = "168h" # Maximum age (7 days)
Consumer Settings
Tune consumer for your workload:
[nats.consumer_rules]
max_ack_pending = 1024 # Increase for higher throughput
ack_wait = "30s" # Adjust based on processing time
backoff = ["5s", "30s", "2m", "5m"] # Retry delays
Security Considerations
-
Use SSL/TLS for all connections:
- PostgreSQL:
sslmode=require - Telemetry:
insecure = false
- PostgreSQL:
-
Secure secrets - Use environment variables or secret management:
- Database passwords
- NATS credentials
- API keys
-
Network security:
- Use private networks for internal services
- Restrict access to monitoring endpoints
- Use firewall rules appropriately
-
Logging - Avoid logging sensitive data:
- Don't log message payloads in production
- Use appropriate log levels
Backup and Recovery
Database Backups
Ensure regular backups of PostgreSQL/TimescaleDB:
- Use pg_dump or TimescaleDB backup tools
- Test restore procedures regularly
JetStream State
JetStream state is stored in NATS:
- Ensure NATS cluster has proper backup procedures
- Stream data is replicated across cluster nodes
- Test disaster recovery procedures
Message Replay
If needed, messages can be replayed from JetStream:
# Replay from a specific sequence
./bin/receiver listen --replay-from seq:12345
# Replay from a specific time
./bin/receiver listen --replay-from time:2024-11-15T08:00:00Z
Support
For issues or questions:
- Check logs:
journalctl -u caatsm(systemd) or container logs - Review metrics in Prometheus/Grafana
- Check health endpoints:
curl http://localhost:2112/readyz - Consult deployment-specific documentation