Terraform on EKS¶
Use this path to stand up a greenfield AWS deployment. The module in
deploy/terraform/aws/ provisions everything the Helm chart references but does
not manage itself, then emits Helm values that match the chart's schema. It is
the reference deployment the AWS validation rounds run against.
What it creates¶
| Resource | Notes |
|---|---|
| VPC | 10.42.0.0/16, three availability zones, a single NAT gateway. Optional: set create_vpc=false to use an existing VPC. |
| EKS cluster | Managed control plane plus a managed node group, at the version pinned in variables.tf. |
| RDS Postgres | The Iceberg catalog. db.t4g.micro by default, which is a benchmark size. Raise rds_instance_class for production. |
| S3 bucket | The warehouse, with versioning, AES256 encryption at rest, and an optional Glacier-IR lifecycle rule. |
| EFS | The ReadWriteMany (RWX) WAL volume: a file system, a mount target per private subnet, an NFS security group, IRSA for the CSI controller, and the EFS CSI driver as an EKS managed addon. |
| IRSA role | Read and write on the warehouse bucket, trusted by the cluster OIDC provider, scoped to the siglake/siglake ServiceAccount. |
| ECR repo | For one-off image pushes, plus an optional pull-through cache against ghcr.io. |
| Secrets Manager secret | RDS credentials in the JSON shape the chart expects: host, port, user, password, database. |
IRSA is IAM Roles for Service Accounts, the EKS mechanism that maps a Kubernetes ServiceAccount onto an IAM role.
Siglake's supported power-loss guarantee covers ext4 or xfs on a node-attached
volume. On this EFS path, durability follows EFS's NFS fsync(2) and rename
semantics.
Prerequisites¶
- A Siglake checkout at the version you want to install.
- Terraform 1.6 or later.
- AWS CLI 2.13 or later, authenticated, with permissions for VPC, EKS, RDS, IAM, S3, EFS and ECR.
- Helm 3.
kubectlcompatible with the EKS version set indeploy/terraform/aws/variables.tf.- Python 3 and
curlfor Secret creation and the smoke test. - Unrestricted egress, so Terraform can fetch registry modules.
Apply¶
- Run the procedure in one shell. Start at the root of your Siglake checkout
and record that directory before changing into the Terraform module. If any
command fails or a required output is empty, the block returns nonzero and
clears the setup outputs. It publishes
VALUES_FILEonly after Terraform writes a nonempty file and the block returns to the checkout root.
terraform_output_is_set() {
if [ -z "$2" ]; then
printf 'terraform output %s is empty\n' "$1" >&2
return 1
fi
}
terraform_values_are_ready() {
if [ ! -s "$1" ]; then
echo 'terraform output helm_values is empty' >&2
return 1
fi
}
CHECKOUT_ROOT=
REGION=
CLUSTER=
SECRET_ARN=
EFS_ID=
VALUES_FILE=
TERRAFORM_VALUES_FILE=
terraform_bootstrap() {
CHECKOUT_ROOT="$(pwd)" &&
NAME=siglake-prod &&
APPLY_REGION=us-west-2 &&
NAMESPACE=siglake &&
RELEASE=siglake &&
cd "$CHECKOUT_ROOT/deploy/terraform/aws" &&
terraform init &&
terraform apply \
-var "name=$NAME" \
-var "region=$APPLY_REGION" &&
REGION="$(terraform output -raw region)" &&
terraform_output_is_set region "$REGION" &&
CLUSTER="$(terraform output -raw cluster_name)" &&
terraform_output_is_set cluster_name "$CLUSTER" &&
SECRET_ARN="$(terraform output -raw rds_secret_arn)" &&
terraform_output_is_set rds_secret_arn "$SECRET_ARN" &&
EFS_ID="$(terraform output -raw efs_file_system_id)" &&
terraform_output_is_set efs_file_system_id "$EFS_ID" &&
TERRAFORM_VALUES_FILE="$(mktemp -t siglake-helm-values.XXXXXX.yaml)" &&
terraform output -raw helm_values > "$TERRAFORM_VALUES_FILE" &&
terraform_values_are_ready "$TERRAFORM_VALUES_FILE" &&
cd "$CHECKOUT_ROOT" &&
VALUES_FILE="$TERRAFORM_VALUES_FILE"
}
if terraform_bootstrap; then
TERRAFORM_VALUES_FILE=
else
BOOTSTRAP_STATUS=$?
if [ -n "$TERRAFORM_VALUES_FILE" ]; then
rm -f "$TERRAFORM_VALUES_FILE" || :
fi
REGION=
CLUSTER=
SECRET_ARN=
EFS_ID=
VALUES_FILE=
if [ -n "$CHECKOUT_ROOT" ]; then
cd "$CHECKOUT_ROOT" 2>/dev/null || :
fi
(exit "$BOOTSTRAP_STATUS")
fi
The full plan creates roughly 60 resources and takes about 15 minutes, most of it the EKS control plane.
Post-apply¶
- Write the cluster context to a fresh kubeconfig, confirm that
kubectlreaches that cluster, then create the release namespace. Themktempfile starts empty, so no context already in your kubeconfig can answer for the new cluster. Each command runs only if the one before it succeeded: if the chain stops, nothing is created and$?is nonzero.
KUBECONTEXT="$CLUSTER"
KUBECONFIG="$(mktemp -t siglake-kubeconfig.XXXXXX)" &&
export KUBECONFIG &&
printf 'apiVersion: v1\nkind: Config\n' > "$KUBECONFIG" &&
aws eks update-kubeconfig \
--region "$REGION" \
--name "$CLUSTER" \
--alias "$KUBECONTEXT" \
--kubeconfig "$KUBECONFIG" &&
SELECTED_CONTEXT="$(kubectl config current-context)" &&
if [ "$SELECTED_CONTEXT" != "$KUBECONTEXT" ]; then
echo "kubeconfig selects $SELECTED_CONTEXT, not $KUBECONTEXT" >&2
false
fi &&
kubectl --context "$KUBECONTEXT" get nodes &&
NAMESPACE_MANIFEST="$(kubectl --context "$KUBECONTEXT" create namespace \
"$NAMESPACE" --dry-run=client -o yaml)" &&
printf '%s\n' "$NAMESPACE_MANIFEST" \
| kubectl --context "$KUBECONTEXT" apply -f -
$KUBECONTEXT and $KUBECONFIG are what the rest of the procedure targets.
Steps 3 to 7 name that context on every kubectl and helm call. The smoke
test calls kubectl without a context argument, so it follows the current
context in the exported KUBECONFIG, which is the context this step
checked.
- Create the EFS StorageClass. The Terraform module installs the EFS CSI driver, but it does not create Kubernetes objects.
cat <<YAML | kubectl --context "$KUBECONTEXT" apply -f -
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: efs-sc
provisioner: efs.csi.aws.com
parameters:
provisioningMode: efs-ap
fileSystemId: $EFS_ID
directoryPerms: "700"
reclaimPolicy: Delete
volumeBindingMode: Immediate
YAML
- Copy the RDS credentials from AWS Secrets Manager into the namespace. The
Secret name matches
postgres.existingSecretin the generated Helm values. Write the converter to a temporary file first, because the secret JSON takes standard input. The JSON stays off every command line, and the converter prints a manifest only after it finds usable values for all five fields. A failed retrieval or malformed payload sends nothing tokubectl.
(
SECRET_SCRIPT=""
SECRET_JSON=""
SECRET_MANIFEST=""
cleanup_secret() {
SECRET_STATUS=$?
trap - 0 HUP INT TERM
unset SECRET_JSON SECRET_MANIFEST
if [ -n "$SECRET_SCRIPT" ] && ! rm -f "$SECRET_SCRIPT"; then
[ "$SECRET_STATUS" -ne 0 ] || SECRET_STATUS=1
fi
exit "$SECRET_STATUS"
}
trap cleanup_secret 0
trap 'exit 1' HUP INT TERM
SECRET_SCRIPT="$(mktemp -t siglake-secret.XXXXXX.py)" || exit
cat > "$SECRET_SCRIPT" <<'PY' || exit
import base64
import json
import sys
REQUIRED = ("host", "port", "user", "password", "database")
namespace, name = sys.argv[1:3]
try:
payload = json.loads(sys.stdin.read())
except ValueError:
sys.exit("secret payload is not valid JSON")
if not isinstance(payload, dict):
sys.exit("secret payload is not a JSON object")
missing = [
key
for key in REQUIRED
if key not in payload
or (isinstance(payload[key], str) and not payload[key].strip())
]
if missing:
sys.exit(f"secret payload is missing: {', '.join(missing)}")
unusable = []
for key in REQUIRED:
value = payload[key]
if key == "port":
valid_type = isinstance(value, str) or (
isinstance(value, int) and not isinstance(value, bool)
)
else:
valid_type = isinstance(value, str)
if not valid_type:
unusable.append(key)
if unusable:
sys.exit(f"secret payload has unusable fields: {', '.join(unusable)}")
lines = [
"apiVersion: v1",
"kind: Secret",
"metadata:",
f" name: {name}",
f" namespace: {namespace}",
"type: Opaque",
"data:",
]
for key in REQUIRED:
encoded = base64.b64encode(str(payload[key]).encode()).decode()
lines.append(f" {key}: {encoded}")
print("\n".join(lines))
PY
if SECRET_JSON="$(aws secretsmanager get-secret-value \
--region "$REGION" \
--secret-id "$SECRET_ARN" \
--query SecretString \
--output text)"; then
:
else
SECRET_STATUS=$?
echo "Secrets Manager retrieval failed; no Secret applied." >&2
exit "$SECRET_STATUS"
fi
SECRET_MANIFEST="$(
printf '%s' "$SECRET_JSON" \
| python3 "$SECRET_SCRIPT" "$NAMESPACE" "${NAME}-postgres"
)" || exit
printf '%s\n' "$SECRET_MANIFEST" \
| kubectl --context "$KUBECONTEXT" apply -f -
)
- Confirm that the StorageClass and Secret exist before installing the chart.
kubectl --context "$KUBECONTEXT" get storageclass efs-sc
kubectl --context "$KUBECONTEXT" --namespace "$NAMESPACE" \
get secret "${NAME}-postgres"
- Install the chart from the checkout root. The generated values configure Postgres, S3 and IRSA. The explicit override selects the EFS StorageClass for the write-ahead log (WAL). If Helm fails, the block exits with Helm's status and keeps the values file, so you can fix the cause and run this step again. Stop there: do not continue to verification. The values file is deleted only after Helm reports success.
(
if helm upgrade --install "$RELEASE" "$CHECKOUT_ROOT/deploy/helm/siglake" \
--kube-context "$KUBECONTEXT" \
--namespace "$NAMESPACE" \
--values "$VALUES_FILE" \
--set wal.storageClassName=efs-sc \
--wait --timeout 10m; then
rm -f "$VALUES_FILE"
else
HELM_STATUS=$?
printf 'helm install failed; values kept at %s\n' "$VALUES_FILE" >&2
exit "$HELM_STATUS"
fi
)
- Verify the deployment. Wait for the ingester Deployment and query StatefulSet, then run the smoke test from the checkout root.
(
set -e
kubectl --context "$KUBECONTEXT" --namespace "$NAMESPACE" rollout status \
"deployment/${RELEASE}-ingester" --timeout=5m
kubectl --context "$KUBECONTEXT" --namespace "$NAMESPACE" rollout status \
"statefulset/${RELEASE}-query" --timeout=5m
cd "$CHECKOUT_ROOT"
env "SIGLAKE_"RELEASE="$RELEASE" "SIGLAKE_"NAMESPACE="$NAMESPACE" \
KUBECONFIG="$KUBECONFIG" ./deploy/aws/smoke.sh
)
The outputs you wire into chart values:
| Output | Feeds |
|---|---|
serviceaccount_role_arn |
serviceAccount.annotations."eks.amazonaws.com/role-arn" |
warehouse_bucket |
s3.bucket |
region |
s3.region |
efs_file_system_id |
The post-apply StorageClass. |
helm_values sets postgres.existingSecret to <name>-postgres. The command
above creates that Secret from the module's rds_secret_arn output before Helm
starts the pods.
Helper scripts¶
deploy/aws/ wraps the common flows:
Use up.sh instead of the numbered procedure. Running both repeats the Helm
installation and the Kubernetes setup.
| Script | Purpose |
|---|---|
up.sh |
Automated alternative to this procedure: apply Terraform, configure the cluster, create Secrets and StorageClasses, install Helm and wait for readiness. |
down.sh |
Remove the release and application resources. Keep the EKS cluster and its supporting VPC, EFS and ECR resources by default. |
smoke.sh |
End-to-end correctness check. |
query-bench.sh |
Query benchmark suite. |
warehouse-load.sh |
Load a corpus into the warehouse. |
operator-smoke.sh |
Operator path validation. |
deploy/aws/config/ holds ready-made values files: values.smoke.yaml,
values.multi-replica.yaml and values.query-perf.yaml.
Adjust the defaults for production¶
The defaults are sized for validation runs. Change these before you carry production traffic.
Raise rds_instance_class. The catalog is on the commit path for every drain
batch, and commit duration shows up directly in
siglake_iceberg_commit_duration_seconds. Size it against your commit rate.
Add a NAT gateway per availability zone if you need tolerance for the loss of one zone. The module provisions a single NAT gateway, which is a single-zone dependency.
Move EFS to provisioned or elastic throughput for heavy workloads. The module
provisions bursting throughput; under sustained high ingest, burst credits
deplete and WAL writes slow down.
Decide on the S3 lifecycle rule deliberately. Glacier-IR transitions are
available through warehouse_lifecycle_days_to_glacier and off by default.
Transitioning warehouse objects raises query latency on cold data, because
Siglake does not stage data back.
Request an EC2 vCPU quota increase ahead of time. Multi-ingester topologies hit
vCPU service quotas before anything else. Several benchmark rounds failed with
VcpuLimitExceeded even though the arithmetic fitted, because terminating
instances from an earlier teardown still counted against the quota. Leave a
cooldown after a teardown for the same reason.
Teardown¶
Run the cleanup script from the same shell, so it inherits the KUBECONFIG
step 2 exported and removes the release from the cluster you installed it on.
The block deletes the kubeconfig only after the script reports success.
(
cd "$CHECKOUT_ROOT" || exit
if ./deploy/aws/down.sh; then
rm -f "$KUBECONFIG"
else
DOWN_STATUS=$?
printf 'cleanup failed; kubeconfig kept at %s\n' "$KUBECONFIG" >&2
exit "$DOWN_STATUS"
fi
)
If the directory change or down.sh fails, the block exits with that status and
keeps the kubeconfig. Fix the cause, then run the same block again from the same
shell. The retry needs that kubeconfig to reach the cluster, and the script
skips the Helm release, Secret, PVCs and namespace that the first run already
deleted.
down.sh exits with Terraform's own status when the final terraform destroy
fails, and prints down complete only after a destroy that succeeded. A failed
destroy leaves AWS resources up and billing. The script also fails before it
reaches Terraform, for a missing Terraform directory or an unknown cleanup mode,
so read its last error line to tell the two cases apart. The Helm and kubectl
steps ahead of the destroy are best effort on purpose: the script carries on
when they fail, and their failures do not change its exit status.
The default cluster cleanup mode removes the Helm release, namespace, RDS
instance, warehouse bucket and application IAM resources. It keeps the EKS
cluster, VPC, EFS and ECR resources warm for the next run. The script empties
the versioned warehouse bucket in this mode, including its object versions.
Warning
The script's all cleanup mode deletes the complete Terraform stack. If
you also enable warehouse emptying, it deletes every object version first.