Files
go-caatsm/docs/prod-guide.md
T

459 lines
11 KiB
Markdown
Raw Normal View History

# Production Deployment Guide
This guide covers deploying and running the CAATSM application in production environments.
## Overview
Production deployments require:
- **JetStream mode** (mandatory) - for message reliability and persistence
- **Manual Stream/Consumer creation** - not auto-created in production
- **Proper configuration** - using `configs/config.prod.toml`
- **SSL/TLS connections** - for secure database and telemetry connections
- **High availability** - stream replicas set to 3+ for HA
## Prerequisites
Before deploying to production:
1. **Build the binary:**
```bash
make build
# Binary will be at bin/receiver
```
2. **NATS JetStream cluster** - must be running and accessible
3. **PostgreSQL/TimescaleDB** - database must be accessible with SSL
4. **OpenTelemetry collector** (optional) - for observability
5. **Prometheus** (optional) - for metrics scraping
## Configuration
### 1. Production Configuration File
The repository includes a production configuration template at `configs/config.prod.toml`. Copy and customize it:
```bash
cp configs/config.prod.toml configs/config.prod.toml.local
# Edit config.prod.toml.local with your production values
```
**Key configuration sections:**
```toml
[nats]
url = "nats://nats.prod:4222"
mode = "jetstream" # Production MUST use JetStream
stream = "TELEGRAM"
consumer = "telegram-consumer"
[nats.stream_limits]
max_msgs = 1000000 # Adjust based on requirements
max_bytes = 1073741824 # 1GB
max_age = "168h" # 7 days
discard = "old"
storage = "file"
replicas = 3 # Use 3+ for HA in production
[nats.consumer_rules]
max_deliver = 5
ack_wait = "30s"
max_ack_pending = 1024
deliver_policy = "new" # Start from new messages in production
replay_policy = "instant"
backoff = ["5s", "30s", "2m", "5m"]
[subscription]
topic = "telegram.serial"
queue_group = "tele-queue"
[postgres]
url = "postgres://user:password@db.prod:5432/aviation?sslmode=require"
max_conns = 20
min_conns = 5
[app]
batch_size = 100
batch_timeout = "2s"
monitor_interval = "30s"
[log]
level = "info"
format = "json"
output = ["stdout"]
[telemetry]
enabled = true
endpoint = "otel-collector.prod:4318"
insecure = false # Use TLS in production
[monitoring]
disabled = false
addr = ":2112"
enable_metrics = true
enable_health = true
[dlq]
enabled = true
subject = "caatsm.dlq"
```
### 2. Environment Variables
Alternatively, you can use environment variables instead of a config file:
```bash
export GO_ENV=prod
export CAATSM_NATS_URL=nats://nats.prod:4222
export CAATSM_NATS_MODE=jetstream
export CAATSM_POSTGRES_URL=postgres://user:pass@db:5432/aviation?sslmode=require
export CAATSM_LOG_LEVEL=info
export CAATSM_TELEMETRY_ENABLED=true
export CAATSM_TELEMETRY_ENDPOINT=otel-collector.prod:4318
export CAATSM_TELEMETRY_INSECURE=false
```
## JetStream Setup
### Pre-create Stream and Consumer
**IMPORTANT:** The application does NOT auto-create streams/consumers in production. You must create them manually before starting the application.
#### Using NATS CLI
```bash
# Create stream
nats stream add TELEGRAM \
--subjects "telegram.serial,telegram.json" \
--storage file \
--replicas 3 \
--max-msgs 1000000 \
--max-bytes 1GB \
--max-age 7d \
--discard old
# Create consumer
nats consumer add TELEGRAM telegram-consumer \
--filter "telegram.serial" \
--ack explicit \
--deliver new \
--max-deliver 5 \
--ack-wait 30s \
--max-pending 1024
```
#### Using NATS Management API
You can also create streams/consumers programmatically using the NATS management API or configuration files.
### Verify Setup
```bash
# Check stream exists
nats stream info TELEGRAM
# Check consumer exists
nats consumer info TELEGRAM telegram-consumer
# Test connection
nats pub telegram.serial "ZCZC TEST 150631..."
```
## Running the Application
### Using Make
```bash
make run-prod # GO_ENV=prod (requires config.prod.toml)
```
### Using Task
```bash
task run-prod # GO_ENV=prod (requires config.prod.toml)
```
### Direct Execution
```bash
# Using binary with config file
GO_ENV=prod ./bin/receiver listen
# Or with environment variables (no config file needed)
GO_ENV=prod \
CAATSM_NATS_URL=nats://nats.prod:4222 \
CAATSM_NATS_MODE=jetstream \
CAATSM_POSTGRES_URL=postgres://user:pass@db:5432/aviation?sslmode=require \
./bin/receiver listen
```
## Production Checklist
Before deploying to production, verify:
- ✅ `nats.mode = "jetstream"` in config (mandatory)
- ✅ Stream and Consumer created manually
- ✅ Stream replicas set to 3+ for high availability
- ✅ PostgreSQL connection configured with SSL (`sslmode=require`)
- ✅ Log level set to `info` or `warn` (not `debug`)
- ✅ Log format set to `json` for log aggregation
- ✅ Telemetry endpoint configured (if using observability)
- ✅ Telemetry TLS enabled (`insecure = false`)
- ✅ Monitoring endpoints exposed for Prometheus scraping
- ✅ DLQ enabled for poison message handling
- ✅ Appropriate retention limits configured (max_msgs, max_bytes, max_age)
- ✅ Connection pool sizes appropriate for workload
- ✅ Batch sizes tuned for throughput
## Monitoring and Observability
### Health Endpoints
The application exposes health and readiness endpoints:
- `GET /livez` - Liveness endpoint (process health)
- `GET /readyz` - Readiness endpoint (checks PostgreSQL and NATS)
- `GET /healthz` - Alias for `/readyz`
- `GET /metrics` - Prometheus metrics
Configure Prometheus to scrape metrics:
```yaml
scrape_configs:
- job_name: 'caatsm'
static_configs:
- targets: ['caatsm:2112']
```
### Key Metrics
Monitor these metrics in production:
- `caatsm_messages_total{stream,consumer,result}` - Message throughput and results
- `caatsm_handle_latency_seconds_bucket` - Processing latency
- `caatsm_retries_total{stream,consumer,reason}` - Retry counts
- `caatsm_db_queries_total{operation,result}` - Database activity
- `caatsm_dlq_messages_total` - Dead-letter queue messages
- `caatsm_nats_consumer_pending_messages` - Consumer backlog/lag
### Logging
Production logs are in JSON format for easy parsing by log aggregation systems:
```json
{
"level": "info",
"ts": 1234567890.123,
"caller": "nats/consumer.go:123",
"msg": "Started consuming messages",
"subject": "telegram.serial",
"consumer": "telegram-consumer",
"stream": "TELEGRAM"
}
```
## High Availability
### Multiple Instances
Run multiple instances of the application for high availability:
- All instances use the same durable consumer name
- JetStream distributes messages across instances
- Each instance independently fetches messages
- If an instance fails, others continue processing
### Stream Replication
Configure stream with 3+ replicas for HA:
```toml
[nats.stream_limits]
replicas = 3 # Minimum 3 for HA, 5 for better distribution
```
### Database Connection Pooling
Configure appropriate connection pool sizes:
```toml
[postgres]
max_conns = 20 # Adjust based on number of instances
min_conns = 5
```
## Troubleshooting
### Stream Not Found
**Error:** `stream TELEGRAM not found`
**Solution:** Create the stream manually before starting the application (see "JetStream Setup" above).
### Consumer Not Found
**Error:** `consumer telegram-consumer not found in stream TELEGRAM`
**Solution:** Create the consumer manually before starting the application (see "JetStream Setup" above).
### Messages Not Being Consumed
**Symptoms:** High pending count, no messages processed
**Check:**
1. Verify consumer exists: `nats consumer info TELEGRAM telegram-consumer`
2. Check pending messages: `nats consumer next TELEGRAM telegram-consumer`
3. Verify application is running and connected
4. Check logs for errors
**Solutions:**
- Increase `batch_size` if processing is slow
- Add more consumer instances
- Check for processing errors in logs
### High Pending Count
**Symptoms:** Consumer has many pending messages
**Solutions:**
- Increase `batch_size` in config
- Add more application instances
- Check processing latency
- Verify database performance
### Messages Being Redelivered
**Symptoms:** Same messages processed multiple times
**Check:**
- Processing logs for errors
- `ack_wait` timeout may be too short
- Processing may be taking longer than `ack_wait`
**Solutions:**
- Increase `ack_wait` if processing takes longer
- Fix processing errors
- Check database connection and performance
### Connection Issues
**NATS Connection:**
- Verify NATS server is accessible
- Check network connectivity
- Verify NATS URL in config
**PostgreSQL Connection:**
- Verify database is accessible
- Check SSL certificate configuration
- Verify connection string format
- Check firewall rules
## Deployment Options
### Systemd Deployment
See `docs/deploy-systemd.md` for a complete systemd service deployment example.
### Kubernetes Deployment
See `docs/deploy-k8s.md` for Kubernetes deployment with ConfigMap/Secret and health probes.
## Performance Tuning
### Batch Processing
Adjust batch size based on message size and processing time:
```toml
[app]
batch_size = 100 # Increase for higher throughput
batch_timeout = "2s" # Adjust based on latency requirements
```
### Connection Pools
Tune database connection pool:
```toml
[postgres]
max_conns = 20 # Total connections across all instances
min_conns = 5 # Keep-alive connections
```
### Stream Retention
Configure retention based on requirements:
```toml
[nats.stream_limits]
max_msgs = 1000000 # Maximum messages
max_bytes = 1073741824 # Maximum size (1GB)
max_age = "168h" # Maximum age (7 days)
```
### Consumer Settings
Tune consumer for your workload:
```toml
[nats.consumer_rules]
max_ack_pending = 1024 # Increase for higher throughput
ack_wait = "30s" # Adjust based on processing time
backoff = ["5s", "30s", "2m", "5m"] # Retry delays
```
## Security Considerations
1. **Use SSL/TLS** for all connections:
- PostgreSQL: `sslmode=require`
- Telemetry: `insecure = false`
2. **Secure secrets** - Use environment variables or secret management:
- Database passwords
- NATS credentials
- API keys
3. **Network security**:
- Use private networks for internal services
- Restrict access to monitoring endpoints
- Use firewall rules appropriately
4. **Logging** - Avoid logging sensitive data:
- Don't log message payloads in production
- Use appropriate log levels
## Backup and Recovery
### Database Backups
Ensure regular backups of PostgreSQL/TimescaleDB:
- Use pg_dump or TimescaleDB backup tools
- Test restore procedures regularly
### JetStream State
JetStream state is stored in NATS:
- Ensure NATS cluster has proper backup procedures
- Stream data is replicated across cluster nodes
- Test disaster recovery procedures
### Message Replay
If needed, messages can be replayed from JetStream:
```bash
# Replay from a specific sequence
./bin/receiver listen --replay-from seq:12345
# Replay from a specific time
./bin/receiver listen --replay-from time:2024-11-15T08:00:00Z
```
## Support
For issues or questions:
- Check logs: `journalctl -u caatsm` (systemd) or container logs
- Review metrics in Prometheus/Grafana
- Check health endpoints: `curl http://localhost:2112/readyz`
- Consult deployment-specific documentation