Appearance
keel.yml Reference
Reference for every keel.yml field, type, default, and validation rule. Keel reports all validation errors in one run. Optional sections such as database, cache, domain, waf, resources, and environments may be omitted.
keel.yml.example
keel init writes an annotated keel.yml.example alongside your config — every option Keel understands, with notes, kept as a sidecar because Keel's own edits to keel.yml round-trip through the YAML parser. A test suite loads and validates the example, and a reflection test fails if any config field is missing from it, so it cannot drift from this reference.
Top level
| Key | Type | Required | Default | Validation |
|---|---|---|---|---|
name | string | yes | — | RFC 1123 hostname |
region | string | yes | us-east-1 | non-empty |
platform | string | no | — | shared-platform name; an app owns its VPC and cluster when omitted |
services | map | at least one of services/sites | — | names match ^[a-z][a-z0-9_-]*$, no collisions after -→_ normalization |
sites | map | at least one of services/sites | — | see below |
monitoring | section | no | — | see below |
capacity | section | no | Fargate only | see below |
resources | map | no | — | see below |
database | section | no | — | see below |
cache | section | no | — | see below |
vpc | section | no | see below | |
networking | section | no | see below | |
load_balancer | section | no | enabled | enabled, idle_timeout |
domain | section | no | — | see below |
waf | section | no | — | see below |
pipeline | section | yes | — | see below |
release | section | no | — | see below |
custom | section | no | — | directory of user-authored OpenTofu files plus outputs injected as environment variables |
alerts | section | no | — | CloudWatch application alarms (Pro) |
audit | section | no | — | see below (Pro) |
idp | section | no | — | see below (Pro) |
preview | section | no | disabled | see below (Pro) |
environment | map[string]string | no | — | merged into every service; service-level wins |
secrets | list of string | no | — | added to every service (union by name) |
logs.retention_days | int | no | 14 | a CloudWatch-accepted value (1, 3, 5, 7, 14, 30, 60, 90, 120, 150, 180, 365, 400, 545, 731, 1096, 1827, 2192, 2557, 2922, 3288, 3653) or -1 for never |
tags | map[string]string | no | — | applied to every managed AWS resource (ManagedBy=keel is never overridable). Keys that differ only in case — from Keel's App/Environment/ManagedBy tags or from each other — are rejected: IAM treats tag keys case-insensitively |
settings | section | no | both false | verbose, cautious |
environments | map | no | — | see below |
default_environment | string | no | — | required (or resolvable) when multiple environments are declared |
Networking defaults
The former mode preset has been removed. Its three decisions are explicit and independently overridable: load_balancer.enabled defaults to true, networking.public_tasks to false, and networking.nat_gateway to single. The inexpensive arrangement is:
yaml
load_balancer:
enabled: false
networking:
public_tasks: true
nat_gateway: noneservices.<name>
| Key | Type | Required | Default | Validation |
|---|---|---|---|---|
type | enum | yes | — | web | worker | scheduled |
port | int | web only | — | web services must specify a port; rejected on scheduled |
health_check | string or block | web only | — | web services must specify a path; timeout < interval; rejected on scheduled |
dockerfile | string | no | Dockerfile | — |
command | string or list | no | image CMD | argv, not shell; scalar split on whitespace honoring quotes; shell metacharacters (| & ; < > ( ) $ \``) rejected — use ["sh", "-c", "..."]. With buildpacks, names a Procfile process type (["web"]`) |
cpu | int | no | 256 | ≥ 128; Fargate services must use a supported Fargate CPU/memory pairing, while EC2 services must fit the selected instances |
memory | int | no | 512 | > 0; must satisfy the selected fleet's constraints |
desired_count | int | no | 1 | rejected on scheduled; ignored when autoscaling is set |
autoscaling | block | no | — | see below; rejected on scheduled |
visibility | enum | no | public | public | internal; web services only |
auth | block | no | — | idp: true or an oidc block, plus optional unauthenticated_paths; internal web services only |
capacity | block | no | Fargate on demand | fleet: fargate | ec2, spot 0–100, on_demand ≥ 0 |
networking.public_tasks | bool | no | environment default | per-service subnet placement; must agree with an EC2 fleet's placement |
gpu | int | no | — | ≥ 1; requires an EC2 fleet of GPU instances |
stop_timeout | int | no | ECS default, or 120 with Spot | 0–120 seconds |
schedule | block | scheduled only | — | see below; rejected on any other type |
environment | map | no | — | wins over app-level |
secrets | list | no | — | union with app-level |
bindings | list | no | — | see below |
permissions | list | no | — | each needs actions and resources; effect defaults to Allow |
Only one public web service may use the root route. Internal services use service-specific hostnames and can share the internal load balancer.
autoscaling block
| Key | Required | Default | Validation |
|---|---|---|---|
min | yes | — | ≥ 0; 0 requires target_queue so the service can wake up |
max | yes | — | > min |
target_cpu | no | — | 1–100 (% average CPU) |
target_memory | no | — | 1–100 (% average memory) |
target_requests | no | — | ≥ 1; ALB requests per target per minute — web services behind an enabled load balancer only |
target_queue.resource | with queue target | — | a declared sqs resource |
target_queue.backlog_per_task | with queue target | — | ≥ 1 visible messages per running task |
scale_in_cooldown | no | 300 | 0–3600 seconds; explicit 0 is valid |
scale_out_cooldown | no | 60 | 0–3600 seconds |
At least one target is required. When multiple targets are configured, each calculates a desired count and ECS uses the largest. Queue depth is the only signal available while no task is running, so it is also the only target that supports min: 0. Manage settings with keel autoscale show/set/off; set accepts --queue and --backlog-per-task.
schedule block
For type: scheduled services — a task definition run by EventBridge Scheduler, with no ECS service behind it. See the Scheduled Tasks guide.
| Key | Required | Default | Validation |
|---|---|---|---|
expression | yes | — | EventBridge syntax: cron(...) (six fields; ? required for the unused day field; day-of-week 1–7 with 1 = Sunday), rate(n unit) (n ≥ 1), or at(yyyy-mm-ddThh:mm:ss) |
timezone | no | UTC | an IANA zone name, validated against embedded tzdata |
enabled | no | true | false creates the schedule DISABLED (PAUSED in the dashboard) |
retries | no | 3 | 0–185; retries a failed invocation, not a non-zero exit |
dead_letter | no | — | must name a resources entry of type: sqs |
health_check block
| Key | Default | Range |
|---|---|---|
path | / | required for web |
grace_period | 60 | 0+ seconds before ECS may replace a new task |
interval | 30 | 5–300 |
timeout | 5 | 2–120; must be < interval |
healthy_threshold | 3 | 2–10 |
unhealthy_threshold | 3 | 2–10 |
matcher | "200-299" | HTTP code range |
deregistration_delay | 30 | 0–3600 |
sites.<name>
Static sites: a private S3 bucket served through CloudFront. See the Static Sites guide. Names match ^[a-z][a-z0-9-]*$ (no _ — the name becomes part of a bucket name) and must not collide with a service name.
sites and services are independent: declare either, or both. An app that declares no services gets no VPC, cluster, or ALB, and may not then declare a database, cache, or waf (each needs the VPC's subnets). It may still declare a pipeline.source, which buys a CodeBuild site build without buying any of the networking.
| Key | Required | Default | Validation |
|---|---|---|---|
root | yes | — | the directory to upload; no default on purpose |
build | no | — | argv command that produces root; a site of committed files needs none |
build_in | no | codebuild when pipeline.source is set, else local | codebuild | local. codebuild requires pipeline.source, and root must then be inside the repository |
spa | no | false | map 403 and 404 to the index document with a 200 — required for client-side routing |
index_document | no | index.html | — |
error_document | no | index.html | — |
cdn | yes | — | the block must be present: the bucket is private and served only through CloudFront |
cdn.enabled | no | true | false takes the site down (the bucket survives) |
cdn.domain | no | *.cloudfront.net name | must sit inside the top-level domain zone; certificate issued in us-east-1 |
cdn.certificate_issued | no | false | domain.dns: external only; written by keel sites verify |
cdn.price_class | no | PriceClass_100 | PriceClass_100 | PriceClass_200 | PriceClass_All |
cdn.compress | no | true | — |
cdn.default_ttl | no | 86400 | 0–31536000 seconds |
monitoring
| Key | Default | Values |
|---|---|---|
container_insights | enabled | disabled | enabled | enhanced |
enabled gives cluster- and service-level metrics; per-task figures are read from the performance logs via a Logs Insights query (billed per GB scanned). enhanced publishes real per-task metrics (billed per metric per month — grows with task count). disabled means no live utilization anywhere; gauges report unknown.
vpc
| Key | Default | Validation |
|---|---|---|
cidr | 10.0.0.0/16 | — |
availability_zones | 2 | 2 or 3 |
networking
| Key | Default | Validation |
|---|---|---|
public_tasks | false | environment default; individual services may override it |
nat_gateway | single | single | per_az | none |
private_aws_endpoints | false | nat_gateway: none with private tasks requires this to be true |
endpoint_azs | all | all | single |
database
| Key | Required | Default | Validation |
|---|---|---|---|
engine | yes | — | postgres | mysql | aurora-postgres | aurora-mysql |
version | yes | — | non-empty |
instance | yes | — | non-empty |
name | no | keeldb | ^[A-Za-z][A-Za-z0-9_]*$ |
username | no | keeluser | same pattern; not a reserved name (admin, root, postgres, …) |
storage | yes | 20 | ≥ 20 |
multi_az | no | false | pointer: an environment override that omits it inherits the base |
retain | no | true | deletion protection + guaranteed final snapshot |
backup_retention_days | no | 7 | 0–35; 0 requires retain: false |
backup_window | no | AWS chooses | e.g. "04:00-05:00" |
log_exports | no | [] (none) | postgres engines: postgresql, upgrade; mysql engines: error, general, slowquery, audit. Nothing is exported by default — CloudWatch ingestion is billed per GB. Required for keel db logs |
iam_auth | no | false | enables per-person database roles managed by keel db grant/revoke (Pro) |
cache
| Key | Required | Default | Validation |
|---|---|---|---|
engine | yes | — | valkey | redis |
version | no | — | — |
node_type | yes | — | non-empty |
num_nodes | no | 1 | — |
The endpoint is TLS-only; use the injected CACHE_URL / REDIS_URL (rediss://).
domain
See the Custom Domains guide for the three DNS modes.
| Key | Required | Default | Validation |
|---|---|---|---|
name | yes | — | FQDN |
dns | no | route53 | route53 | route53-managed | external |
zone | route53 modes | for route53-managed, the domain itself | FQDN; ignored for external |
zone_id | no | — | pins the zone when a private zone shares its name; not valid with route53-managed |
aliases | no | — | extra names on the certificate; FQDNs or wildcards (*.example.com); every name must sit inside the zone. Wildcards get no alias record — covered by the cert, deliberately not routed |
certificate_arn | no | — | use an existing ACM certificate (must be in the ALB's region); Keel then issues nothing |
certificate_issued | no | false | external only, and only when Keel issues the certificate; written by keel domains verify |
force_https | no | true once TLS is ready | never redirects to a listener that doesn't exist yet — a pending certificate can't take the app offline |
validation_timeout | no | 45m (10m for a Keel-created zone) | positive Go duration, e.g. 20m |
With services, a domain requires a load balancer (the alias record points at the ALB). A sites-only app may declare a domain too — it names the zone that a site's CloudFront alias goes into.
waf
| Key | Default |
|---|---|
enabled | false |
managed_rules | — (CLI default when enabling: AWSManagedRulesCommonRuleSet) |
Requires a load balancer.
capacity
Fargate and Fargate Spot are available on every cluster and need no top-level declaration. capacity.ec2 adds an Auto Scaling group behind an ECS capacity provider; services.<name>.capacity.fleet: ec2 selects it.
| Key | Required | Default | Validation |
|---|---|---|---|
ec2.instance_types | yes | — | one or more compatible instance types; all x86_64 and either all GPU or all non-GPU |
ec2.spot | no | 0 | 0–100 percent above the on-demand base |
ec2.on_demand | no | 0 | ≥ 0 instances always bought on demand |
ec2.min | no | 0 | ≥ 0; zero lets ECS managed scaling take an idle fleet down completely |
ec2.max | yes | — | ≥ 1 |
ec2.public | no | follows networking.public_tasks | places instances, not task ENIs, in public subnets |
ec2.network_mode | no | awsvpc | awsvpc | bridge |
bridge gives tasks the instance's network address and avoids one ENI per task, but every task on the host shares the instance security group. It is refused for a shared-platform tenant because that would collapse network identity across applications. A shared platform may own the EC2 fleet itself; tenants then select fleet: ec2 without declaring top-level capacity.
alerts
The block creates CloudWatch alarms and requires email, sns_topic_arn, or both. Threshold keys are optional: absent uses the listed default and 0 disables that alarm. response_time is off unless set.
| Key | Default |
|---|---|
http_5xx | 1 percent |
unhealthy_hosts | 1 |
response_time | off; positive duration such as 2s |
service_cpu, service_memory | 90 percent |
tasks_running | true |
database_cpu | 90 percent |
database_storage_free | 10 percent |
cache_cpu | 90 percent |
audit
| Key | Default | Validation |
|---|---|---|
exec.record | false | enables cluster-level ECS Exec transcripts |
exec.require_reason | false | makes keel exec --reason mandatory; may be used without recording |
exec.retention_days | logs.retention_days | CloudWatch retention value or -1; requires recording |
exec.kms_key_id | AWS-managed encryption | KMS key ARN or alias/…; requires recording |
Recording belongs to the cluster and captures sessions opened outside Keel too. Transcripts live in /keel/exec/<app>-<env> and survive keel destroy. Use keel audit sessions and keel audit transcript to read them.
idp
This app-scoped IdP is for people signing in to an internal web service. It is separate from the account-level operator IdP described in Operators.
| Key | Default | Validation |
|---|---|---|
enabled | false | requires at least one service with auth.idp: true |
domain | <app>-<environment> | globally unique Cognito hosted-UI prefix |
users | [] | email addresses; Cognito sends temporary passwords |
groups | [] | created and included in the token; Keel does not enforce them |
preview
| Key | Default | Validation |
|---|---|---|
enabled | false | — |
base | default_environment | declared environment |
prefix | pv- | marks names that may be synthesized |
ttl | 72h | positive duration |
load_balancer | disabled for previews | same fields as the top-level block |
networking | public tasks, no NAT | same fields as the top-level block |
database | none | none | shared | own |
desired_count | 1 | ≥ 1 |
A preview is named for its branch; a pull-request number is optional metadata. The older preview.mode preset is gone with top-level mode.
pipeline
| Key | Required | Default | Validation |
|---|---|---|---|
source | when services are declared, or a site sets build_in: codebuild | — | github | codecommit |
repo | github: yes; codecommit: no | — | see accepted forms below |
branch | no | main | — |
connection_arn | no | resolved from SSM | github only; a CodeConnections connection ARN |
build | no | auto | auto | docker | buildpack — auto probes for a Dockerfile; the chosen strategy is always printed (Build: <strategy> — <reason>) |
builder | no | heroku/builder:24 | buildpack path only; must be x86_64 (Fargate does no emulation); rejected alongside build: docker |
pack_version | no | 0.36.4 | buildpack path only; pinned so a build is a function of the commit, not the day |
build_image | no | aws/codebuild/amazonlinux-x86_64-standard:5.0 | an image whose name marks the wrong architecture for the CodeBuild environment (aarch64/arm64 vs x86_64/amd64) is rejected at validation; unmarked custom images are accepted |
Accepted repo forms: owner/name, https://github.com/owner/name[.git], [email protected]:owner/name.git (github); a bare name or the full git-codecommit URL (codecommit). source: codecommit with repo unset means Keel creates and owns a repository named <app>-<env>.
release
Runs once per deploy, before any service takes traffic. Shorthand forms:
yaml
release: bundle exec rails db:migrate
release: ["sh", "-c", "rails db:migrate && rails db:seed"]
release:
command: bundle exec rails db:migrate
service: web
timeout: 20m
environment:
MIGRATION_MODE: online
secrets:
- MIGRATOR_DATABASE_URL| Key | Required | Default | Validation |
|---|---|---|---|
command | yes (when block present) | — | same argv rules as service command |
service | no | the sole web service, or the sole service | must name a service |
timeout | no | 15m | positive Go duration |
environment | no | {} | variables added only to the release task |
secrets | no | [] | secret names injected only into the release task |
resources
Each entry declares exactly one of create (Keel owns it) or existing (Keel binds it). type: s3 | sns | sqs | rds | cache | iam-role. Only s3/sns/sqs can be created. retain defaults to true for managed resources.
create options: versioning (s3), encryption (s3, aws-managed), fifo (sns/sqs), visibility_timeout (sqs, 0–43200). A managed s3 resource name cannot contain _.
existing required fields by type: arn (s3/sns/iam-role); arn + url (sqs); endpoint + port + security_group_id (rds/cache). Optional: identifier, secret_arn, secret_env (must match ^[A-Z_][A-Z0-9_]*$), kms_key_arn.
Bindings
yaml
services:
web:
bindings:
- resource: uploads
access: [list, read, write]
prefix: user-content # S3 only
env: UPLOADS # override the injected variable baseCapabilities per type: s3 list, read, write, delete; sns publish; sqs send, consume; rds/cache connect; iam-role assume.
environments
Each entry overrides the base config; every field optional. Overridable: region, account, platform, load_balancer, networking, capacity, services, sites, resources, database, cache, domain, release, audit, environment, secrets, and tags. Declaring account opts the environment into account isolation. A site's cdn block merges field by field; certificate_issued is cleared when an environment overrides cdn.domain to a different hostname.
Merge rules: database/cache/domain and services.<name> merge field by field; environment and tags merge per key (override wins); secrets union by name; bindings merge by resource; permissions append; resources merge per name; release is replaced whole.
Selection precedence: --env → KEEL_ENV → keel env use → default_environment. With no environments: block, a single synthetic default environment is used.
Derived names
| Concept | Value |
|---|---|
| Name prefix | <app>-<env> |
| ECS cluster | <app>-<env>-cluster |
| ECS service | <app>-<env>-<service> |
| Log group | /ecs/<app>-<env> |
| Config vars | /keel/<app>/<env>/<key> |
| Database identifier | <app>-<env>-db |
| Cache replication group | <app>-<env>-cache |
| ECR repository | <app>-<env> (shared; images tagged <service>-<sha>) |
| CodeBuild project | <app>-<env>-build (per Dockerfile) |
| OpenTofu state key | <app>/<env>/terraform.tfstate |
Worked example
The Rails web+worker config that docs/deploying-rails.md walks through lives at internal/config/testdata/rails-web-worker.yml and is loaded, validated, and HCL-generated by the test suite — the docs cannot drift from it. See Deploying Rails for the annotated version.