Skip to content

Latest commit

Β 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

AgentGuard Infrastructure

Production-ready, cost-optimized Amazon EKS cluster infrastructure using Terraform. Designed to maintain monthly costs under $110 while providing essential Kubernetes functionality for the AgentGuard service.

πŸ—οΈ Architecture Overview

This infrastructure provisions a minimal-cost EKS cluster with the following components:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        AWS Cloud                             β”‚
β”‚                                                              β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚              VPC (10.0.0.0/16)                         β”‚ β”‚
β”‚  β”‚                                                        β”‚ β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”‚ β”‚
β”‚  β”‚  β”‚  AZ us-east-1a   β”‚    β”‚  AZ us-east-1b   β”‚        β”‚ β”‚
β”‚  β”‚  β”‚                  β”‚    β”‚                  β”‚        β”‚ β”‚
β”‚  β”‚  β”‚  Public Subnet   β”‚    β”‚  Public Subnet   β”‚        β”‚ β”‚
β”‚  β”‚  β”‚  10.0.1.0/28     β”‚    β”‚  10.0.2.0/28     β”‚        β”‚ β”‚
β”‚  β”‚  β”‚  (14 IPs)        β”‚    β”‚  (14 IPs)        β”‚        β”‚ β”‚
β”‚  β”‚  β”‚                  β”‚    β”‚                  β”‚        β”‚ β”‚
β”‚  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚    β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚        β”‚ β”‚
β”‚  β”‚  β”‚  β”‚ t2.micro   β”‚  β”‚    β”‚  β”‚ t2.micro   β”‚  β”‚        β”‚ β”‚
β”‚  β”‚  β”‚  β”‚ Worker     β”‚  β”‚    β”‚  β”‚ Worker     β”‚  β”‚        β”‚ β”‚
β”‚  β”‚  β”‚  β”‚ Node       β”‚  β”‚    β”‚  β”‚ Node       β”‚  β”‚        β”‚ β”‚
β”‚  β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚    β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚        β”‚ β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β”‚ β”‚
β”‚  β”‚           β”‚                       β”‚                   β”‚ β”‚
β”‚  β”‚           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                   β”‚ β”‚
β”‚  β”‚                       β”‚                               β”‚ β”‚
β”‚  β”‚              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”                      β”‚ β”‚
β”‚  β”‚              β”‚ Internet Gateway β”‚                      β”‚ β”‚
β”‚  β”‚              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                      β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                                                              β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚           EKS Control Plane (Managed by AWS)           β”‚ β”‚
β”‚  β”‚                  Kubernetes 1.28                       β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                                                              β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚  S3 Bucket   β”‚         β”‚  DynamoDB Table              β”‚ β”‚
β”‚  β”‚  Terraform   β”‚         β”‚  State Locking               β”‚ β”‚
β”‚  β”‚  State       β”‚         β”‚                              β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Key Features

  • βœ… Cost-Optimized: ~$88-110/month total infrastructure cost
  • βœ… High Availability: Multi-AZ deployment across 2 availability zones
  • βœ… Managed Control Plane: AWS-managed EKS control plane (no maintenance overhead)
  • βœ… Auto-Scaling: Configurable node group scaling (1-3 nodes)
  • βœ… Secure State Management: S3 backend with DynamoDB locking
  • βœ… Production-Ready: Encryption at rest, versioned state, IAM best practices

Architecture Decisions

Public Subnet Architecture: Worker nodes are deployed in public subnets with direct internet access to eliminate NAT Gateway costs (~$32/month savings). This is suitable for development and testing environments. For production, consider private subnets with NAT Gateway or VPN access.

t2.micro Instances: Cost-optimized instance type (1 vCPU, 1GB RAM) suitable for lightweight workloads. Supports ~10-15 pods per node.

/28 Subnets: Minimal IP allocation (14 usable IPs per subnet) to reduce costs. Limits node scaling to ~6-7 nodes per subnet.

πŸ’° Cost Breakdown

Component Configuration Monthly Cost
EKS Control Plane Managed by AWS $73.00
Worker Nodes 2Γ— t2.micro instances $15.00
EBS Volumes 2Γ— 20GB gp3 volumes $3.20
Load Balancer Classic ELB + data transfer ~$16.20
Total $88-107

Cost Optimization Tips

  • Scale down during off-hours: Reduce to 1 node to save $7.50/month
  • Use Spot instances: Up to 90% savings for non-critical workloads
  • Delete unused LoadBalancers: Save $16.20/month per LoadBalancer
  • Optimize pod resources: Maximize node utilization to avoid unnecessary scaling

Cost Increase Factors

  • NAT Gateway: +$32/month (not used in this architecture)
  • Larger instances (t2.small): +$7.50/month per node
  • Additional EBS storage: +$0.08/GB/month
  • CloudWatch Logs: +$0.50/GB ingested
  • Data transfer: +$0.09/GB egress

πŸ“‹ Prerequisites

Required Tools

Tool Version Purpose
Terraform >= 1.5.0 Infrastructure provisioning
AWS CLI >= 2.0 AWS authentication and resource management
kubectl >= 1.28 Kubernetes cluster interaction

Verify installations:

terraform --version
aws --version
kubectl version --client

AWS Account Requirements

IAM Permissions

Your AWS credentials must have permissions for:

  • Amazon EKS: Cluster and node group management
  • Amazon EC2: VPC, subnet, security group, instance management
  • Amazon S3: State bucket management
  • Amazon DynamoDB: State lock table management
  • AWS IAM: Role creation and policy attachment
  • Elastic Load Balancing: LoadBalancer provisioning

Service Quotas

Verify your account has sufficient quotas:

  • EKS clusters per region: β‰₯ 1
  • VPCs per region: β‰₯ 1
  • EC2 instances (t2.micro): β‰₯ 2-3 on-demand instances
  • No Elastic IPs required

Cost Budget

Ensure you have approval for ~$100/month infrastructure costs.

πŸš€ Quick Start

Step 1: Configure AWS Credentials

aws configure

Enter your AWS Access Key ID, Secret Access Key, default region, and output format.

Verify:

aws sts get-caller-identity

Step 2: Create Backend Resources

Terraform requires an S3 bucket and DynamoDB table for state management. Choose one method:

Option A: Automated Script (Recommended)

Linux/Mac:

cd agentguard-infrastructure
chmod +x setup-backend.sh
./setup-backend.sh

Windows (PowerShell):

cd agentguard-infrastructure
.\setup-backend.ps1

Option B: Manual Setup

Follow the detailed instructions in MANUAL_SETUP_INSTRUCTIONS.md

What gets created:

  • S3 bucket: agentguard-terraform-state-<random-suffix> (versioning + encryption enabled)
  • DynamoDB table: agentguard-terraform-state-lock (state locking)

See BACKEND_SETUP_README.md for more details.

Step 3: Configure Variables

Copy the example configuration:

cp terraform.tfvars.example terraform.tfvars

Edit terraform.tfvars with your values:

aws_region          = "us-east-1"
cluster_name        = "agentguard-cluster"  # Must be unique in your region
cluster_version     = "1.28"
vpc_cidr            = "10.0.0.0/16"
availability_zones  = ["us-east-1a", "us-east-1b"]
public_subnet_cidrs = ["10.0.1.0/28", "10.0.2.0/28"]
node_instance_types = ["t2.micro"]
node_desired_size   = 2
node_min_size       = 1
node_max_size       = 3
node_disk_size      = 20

Step 4: Update Backend Configuration

Edit backend.tf with your S3 bucket name and region:

terraform {
  backend "s3" {
    bucket         = "agentguard-terraform-state-<your-suffix>"
    key            = "agentguard/terraform.tfstate"
    region         = "us-east-1"
    dynamodb_table = "agentguard-terraform-state-lock"
    encrypt        = true
  }
}

Step 5: Deploy Infrastructure

# Initialize Terraform
terraform init

# Validate configuration
terraform validate
terraform fmt -check

# Preview changes
terraform plan

# Deploy infrastructure (takes ~15-20 minutes)
terraform apply

Type yes when prompted to confirm.

Step 6: Configure kubectl

# Get the kubeconfig command from Terraform output
terraform output kubeconfig_command

# Run the command (example):
aws eks update-kubeconfig --region us-east-1 --name agentguard-cluster

# Verify connectivity
kubectl get nodes

You should see 2 nodes in "Ready" state.

Step 7: Verify Deployment

# Check cluster info
kubectl cluster-info

# View validation deployment
kubectl get deployments
kubectl get pods

# Get LoadBalancer endpoint
terraform output loadbalancer_endpoint

# Test endpoint (wait a few minutes for ELB provisioning)
curl -I http://<loadbalancer-endpoint>

Expected: HTTP 200 response from nginx validation service.

πŸ“ Project Structure

agentguard-infrastructure/
β”œβ”€β”€ main.tf                    # Root module - integrates all components
β”œβ”€β”€ variables.tf               # Input variables
β”œβ”€β”€ outputs.tf                 # Output values
β”œβ”€β”€ versions.tf                # Provider version constraints
β”œβ”€β”€ backend.tf                 # S3 backend configuration
β”œβ”€β”€ terraform.tfvars.example   # Example variable values
β”œβ”€β”€ setup-backend.sh           # Automated backend setup (Linux/Mac)
β”œβ”€β”€ setup-backend.ps1          # Automated backend setup (Windows)
β”œβ”€β”€ README.md                  # This file
β”œβ”€β”€ SETUP.md                   # Detailed setup guide
β”œβ”€β”€ OPERATIONS.md              # Operational procedures
β”œβ”€β”€ TROUBLESHOOTING.md         # Common issues and solutions
β”œβ”€β”€ BACKEND_SETUP_README.md    # Backend setup options
β”œβ”€β”€ MANUAL_SETUP_INSTRUCTIONS.md  # Manual AWS Console setup
└── modules/
    β”œβ”€β”€ vpc/                   # VPC networking module
    β”‚   β”œβ”€β”€ main.tf
    β”‚   β”œβ”€β”€ variables.tf
    β”‚   β”œβ”€β”€ outputs.tf
    β”‚   └── README.md
    β”œβ”€β”€ eks/                   # EKS cluster module
    β”‚   β”œβ”€β”€ main.tf
    β”‚   β”œβ”€β”€ variables.tf
    β”‚   β”œβ”€β”€ outputs.tf
    β”‚   └── README.md
    └── node-group/            # Managed node group module
        β”œβ”€β”€ main.tf
        β”œβ”€β”€ variables.tf
        β”œβ”€β”€ outputs.tf
        └── README.md

πŸ“š Documentation

πŸ”§ Operations

Scaling Worker Nodes

Update terraform.tfvars:

node_desired_size = 3  # Scale to 3 nodes

Apply changes:

terraform apply

Upgrading Kubernetes Version

Update terraform.tfvars:

cluster_version = "1.29"  # Upgrade to 1.29

Apply changes:

terraform apply

Note: Test upgrades in a non-production environment first.

Viewing Terraform Outputs

# View all outputs
terraform output

# View specific output
terraform output cluster_endpoint
terraform output loadbalancer_endpoint

Destroying Infrastructure

terraform destroy

Type yes to confirm. This will delete all resources except the S3 bucket and DynamoDB table.

See OPERATIONS.md for detailed operational procedures.

πŸ› Troubleshooting

Common Issues

Issue Cause Solution
"Error acquiring the state lock" Another Terraform process holds the lock Wait for completion or terraform force-unlock <lock-id>
"InsufficientFreeAddressesInSubnet" Node scaling exceeds /28 subnet capacity Reduce node_max_size or expand subnet CIDR
Nodes not joining cluster Security group, IAM, or subnet issues Verify security groups, IAM policies, and subnet tags
LoadBalancer stuck in "Pending" Missing subnet tags or IAM permissions Check kubernetes.io/role/elb=1 tag and IAM permissions

See TROUBLESHOOTING.md for detailed solutions.

πŸ”’ Security Considerations

Current Architecture (Cost-Optimized)

  • βœ… Public subnets with direct internet access (saves ~$32/month)
  • βœ… Public EKS endpoint access
  • βœ… Encryption at rest for EBS volumes
  • βœ… S3 state encryption enabled
  • βœ… IAM roles follow least privilege principle

Production Recommendations

For production deployments, consider:

  • Private Subnets: Use private subnets with NAT Gateway for worker nodes
  • VPN/Bastion: Implement VPN or bastion host for secure cluster access
  • Network Policies: Enable Kubernetes network policies for pod-to-pod traffic control
  • Audit Logging: Enable EKS audit logging to CloudWatch Logs
  • Pod Security: Implement Pod Security Standards (Restricted profile)
  • Secrets Management: Use AWS Secrets Manager or external secrets operator
  • Private Endpoint: Enable private EKS endpoint access

πŸ§ͺ Testing

Validation Tests

# Terraform validation
terraform validate
terraform fmt -check

# Cluster connectivity
kubectl get nodes
kubectl get pods --all-namespaces

# LoadBalancer test
curl -I http://$(terraform output -raw loadbalancer_endpoint)

# Pod deployment test
kubectl run test-nginx --image=nginx --restart=Never
kubectl wait --for=condition=Ready pod/test-nginx --timeout=60s
kubectl delete pod test-nginx

Property-Based Tests

The infrastructure includes Terratest-based property tests:

  • Cost optimization invariant (total cost < $110)
  • High availability invariant (resources in β‰₯ 2 AZs)
  • Network connectivity invariant (nodes reach internet)
  • Security group invariant (required ports open)
  • State consistency invariant (no drift after apply)

See .kiro/specs/agentguard-infrastructure/design.md for testing strategy details.

🀝 Contributing

Code Standards

  • Run terraform fmt before committing
  • Run terraform validate to check syntax
  • Update documentation for any changes
  • Test changes in a separate environment

Module Development

Each module should include:

  • main.tf: Resource definitions
  • variables.tf: Input variables with descriptions
  • outputs.tf: Output values with descriptions
  • README.md: Module documentation

πŸ“Š Monitoring

CloudWatch Metrics

Monitor these key metrics:

  • EKS Cluster: API server response time, request rate
  • Worker Nodes: CPU utilization, memory utilization, disk usage
  • Pods: Pod count, restart count, resource usage

Cost Monitoring

# View current month costs
aws ce get-cost-and-usage \
    --time-period Start=$(date +%Y-%m-01),End=$(date +%Y-%m-%d) \
    --granularity MONTHLY \
    --metrics BlendedCost \
    --group-by Type=SERVICE

Set up AWS Budgets to receive alerts when costs exceed thresholds.

πŸ”„ Upgrade Path

Kubernetes Version Upgrades

  1. Review EKS version release notes
  2. Test upgrade in non-production environment
  3. Update cluster_version in terraform.tfvars
  4. Run terraform apply
  5. Update node group AMI if needed
  6. Verify workload compatibility

Terraform Version Upgrades

  1. Review Terraform upgrade guide
  2. Update required_version in versions.tf
  3. Run terraform init -upgrade
  4. Test with terraform plan
  5. Apply changes

πŸ“ž Support

Resources

Getting Help

  1. Check TROUBLESHOOTING.md for common issues
  2. Review Terraform plan output for error messages
  3. Check AWS CloudWatch Logs for cluster and node logs
  4. Review the design document for architecture details

πŸ“ License

This infrastructure code is provided as-is for the AgentGuard project.

🎯 Next Steps

After successful deployment:

  1. Deploy Applications: Use kubectl or Helm to deploy your applications
  2. Configure Monitoring: Set up CloudWatch metrics and alarms
  3. Implement Autoscaling: Configure HPA and Cluster Autoscaler
  4. Set Up CI/CD: Integrate with your CI/CD pipeline
  5. Review Security: Implement additional security hardening
  6. Cost Optimization: Monitor and optimize resource usage

Estimated Deployment Time: 15-20 minutes
Estimated Monthly Cost: $88-110
Maintenance Overhead: Low (AWS-managed control plane)

About

AgentGurad Infra

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages