# Configuration (keel.yml)

`keel.yml` describes the application and infrastructure Keel manages. This guide covers the main sections; the [keel.yml reference](../reference/keel-yml) lists every field, type, default, and validation rule.

## A complete example

```yaml
# keel.yml
name: my-api
region: us-east-1

networking:
  public_tasks: false     # private subnets by default
  nat_gateway: single    # single | per_az | none
load_balancer:
  enabled: true

services:
  web:
    type: web             # web | worker | scheduled
    port: 8080
    health_check: /health # or a block, see below
    command: ["bundle", "exec", "puma"]   # overrides the image CMD
    cpu: 256              # 256 | 512 | 1024 | 2048 | 4096
    memory: 512
    desired_count: 2
    # networking:
    #   public_tasks: true # override placement for only this service
    # capacity:
    #   spot: 50           # use Fargate Spot above the on-demand base
    #   on_demand: 1
    autoscaling:
      min: 2
      max: 10
      target_cpu: 60

environment:              # merged into every service
  RAILS_ENV: production
secrets:                  # added to every service, read from SSM
  - SECRET_KEY_BASE

release:                  # runs once per deploy, before any service takes traffic
  command: bundle exec rails db:migrate

logs:
  retention_days: 365     # default 14; -1 never expires

database:
  engine: postgres        # postgres | mysql | aurora-postgres | aurora-mysql
  version: "16"
  instance: db.t4g.micro
  storage: 20
  multi_az: true
  retain: true            # deletion protection + final snapshot (default)
  backup_retention_days: 7
  name: appdb             # default keeldb
  username: appuser       # default keeluser
  # iam_auth: true        # Pro — per-person database identities

cache:
  engine: valkey          # valkey | redis
  node_type: cache.t4g.micro

domain:
  name: api.example.com
  zone: example.com

waf:
  enabled: true

# An optional EC2 fleet beside Fargate, for GPUs or sustained workloads.
# capacity:
#   ec2:
#     instance_types: [g5.xlarge, g5.2xlarge]
#     max: 8

pipeline:
  source: github          # github | codecommit
  repo: myorg/my-api      # required for github; optional for codecommit
  branch: main
  # connection_arn:       # override the org's CodeConnections connection

tags:                     # applied to every managed AWS resource
  team: platform
  cost-center: eng

settings:                 # project-level CLI behavior
  verbose: false
  cautious: false
```

Optional sections (`database`, `cache`, `domain`, `waf`, `resources`, `capacity`, `custom`, `alerts`, `audit`, `idp`, `preview`, and `environments`) can be omitted entirely.

## Services

Each entry under `services:` becomes an ECS workload. Fargate is the default; a service can instead select an EC2 capacity provider declared by the app or its shared platform. Three types exist:

- `web` receives load-balancer traffic and requires a `port` and health check. One public service can own the root route; internal services use service-specific hostnames and can share the internal load balancer.
- `worker` runs continuously without a load balancer or port.
- `scheduled` uses EventBridge Scheduler to start one task per event. It requires a `schedule:` block and does not accept `port`, `health_check`, `desired_count`, or `autoscaling`. See [Scheduled Tasks](./scheduled-tasks).

Static sites are not services — they're a top-level `sites:` block. See [Static Sites](./static-sites).

### Placement and capacity

Placement is per service. `networking.public_tasks` sets the environment default, while `services.<name>.networking.public_tasks` can move one high-egress worker to public subnets without moving a web tier that holds database credentials. Public placement does not open an ingress port; security groups still decide reachability.

`services.<name>.capacity` chooses Fargate or EC2 and the interruptible share. Fargate Spot needs no top-level block. An EC2 service needs `capacity.ec2` on an app-owned cluster, or an EC2 fleet on the shared platform:

```yaml
capacity:
  ec2:
    instance_types: [g5.xlarge, g5.2xlarge]
    spot: 50
    on_demand: 1
    min: 0
    max: 8

services:
  gpu-worker:
    type: worker
    cpu: 4096
    memory: 16384
    gpu: 1
    capacity:
      fleet: ec2
```

One cluster may hold Fargate, Fargate Spot, and an EC2 capacity provider together. Spot services receive a 120-second stop timeout unless they set one explicitly. See the [keel.yml reference](../reference/keel-yml#capacity) for EC2 networking and shared-platform constraints.

### `command`

`command` overrides the image's `CMD`, allowing one image to run different web and worker processes. Without an override, every service runs the image's default command.

Prefer a list of arguments. Keel splits the scalar form on whitespace and rejects shell syntax such as `&&`, `|`, `$`, and redirection. If you need shell behavior, invoke it explicitly: `["sh", "-c", "first && second"]`.

### Health checks

`health_check` accepts a path, or a block when the defaults do not fit:

```yaml
services:
  web:
    health_check:
      path: /up
      grace_period: 90      # seconds before ECS may replace a new task (default 60)
      interval: 30
      timeout: 5
      healthy_threshold: 3
      unhealthy_threshold: 3
      matcher: "200-299"
      deregistration_delay: 30
```

**Set enough startup time**
Keel defaults the health-check grace period to 60 seconds. Increase it when your app takes longer to start; otherwise ECS may replace new tasks before they become healthy. Keel also enables a deployment circuit breaker so failed deployments roll back instead of restarting indefinitely.

## Shared environment and secrets

Anything shared by every service belongs at the top level rather than repeated:

```yaml
environment:            # merged into every service
  RAILS_ENV: production
  APP_HOST: app.example.com
secrets:                # added to every service
  - SECRET_KEY_BASE

services:
  web:
    environment:
      LOG_LEVEL: debug  # a service-level value wins over the app-level one
```

Precedence, lowest to highest: app base → app environment override → service base → service environment override.

## Settings

Two project-level behaviors change how commands communicate, persisted under `settings:` and shared across every environment. Command-line flags override them for a single invocation.

- **verbose** prints each action as it happens (`→ Triggering CodeBuild — my-api-prod-web-build`).
- **cautious** previews state-changing actions and asks for confirmation.

```bash
keel settings                       # list current values
keel settings set cautious true
keel deploy --verbose               # override for one run
```

## Where to go next

- [Deployments](./deployments) — pipeline sources, the deploy flow, rollback
- [Environments](./environments) — staging/production overrides and how they merge
- [AWS Resources & Runtime Access](./resources) — S3/SNS/SQS bindings, databases, caches
- [keel.yml Reference](../reference/keel-yml) — every field with types and defaults
