Skip to content

Terraform on EKS

Use this path to stand up a greenfield AWS deployment. The module in deploy/terraform/aws/ provisions everything the Helm chart references but does not manage itself, then emits Helm values that match the chart's schema. It is the reference deployment the AWS validation rounds run against.

What it creates

Resource Notes
VPC 10.42.0.0/16, three availability zones, a single NAT gateway. Optional: set create_vpc=false to use an existing VPC.
EKS cluster Managed control plane plus a managed node group, at the version pinned in variables.tf.
RDS Postgres The Iceberg catalog. db.t4g.micro by default, which is a benchmark size. Raise rds_instance_class for production.
S3 bucket The warehouse, with versioning, AES256 encryption at rest, and an optional Glacier-IR lifecycle rule.
EFS The ReadWriteMany (RWX) WAL volume: a file system, a mount target per private subnet, an NFS security group, IRSA for the CSI controller, and the EFS CSI driver as an EKS managed addon.
IRSA role Read and write on the warehouse bucket, trusted by the cluster OIDC provider, scoped to the siglake/siglake ServiceAccount.
ECR repo For one-off image pushes, plus an optional pull-through cache against ghcr.io.
Secrets Manager secret RDS credentials in the JSON shape the chart expects: host, port, user, password, database.

IRSA is IAM Roles for Service Accounts, the EKS mechanism that maps a Kubernetes ServiceAccount onto an IAM role.

Siglake's supported power-loss guarantee covers ext4 or xfs on a node-attached volume. On this EFS path, durability follows EFS's NFS fsync(2) and rename semantics.

Prerequisites

  • A Siglake checkout at the version you want to install.
  • Terraform 1.6 or later.
  • AWS CLI 2.13 or later, authenticated, with permissions for VPC, EKS, RDS, IAM, S3, EFS and ECR.
  • Helm 3.
  • kubectl compatible with the EKS version set in deploy/terraform/aws/variables.tf.
  • Python 3 and curl for Secret creation and the smoke test.
  • Unrestricted egress, so Terraform can fetch registry modules.

Apply

  1. Run the procedure in one shell. Start at the root of your Siglake checkout and record that directory before changing into the Terraform module. If any command fails or a required output is empty, the block returns nonzero and clears the setup outputs. It publishes VALUES_FILE only after Terraform writes a nonempty file and the block returns to the checkout root.
terraform_output_is_set() {
  if [ -z "$2" ]; then
    printf 'terraform output %s is empty\n' "$1" >&2
    return 1
  fi
}

terraform_values_are_ready() {
  if [ ! -s "$1" ]; then
    echo 'terraform output helm_values is empty' >&2
    return 1
  fi
}

CHECKOUT_ROOT=
REGION=
CLUSTER=
SECRET_ARN=
EFS_ID=
VALUES_FILE=
TERRAFORM_VALUES_FILE=

terraform_bootstrap() {
  CHECKOUT_ROOT="$(pwd)" &&
  NAME=siglake-prod &&
  APPLY_REGION=us-west-2 &&
  NAMESPACE=siglake &&
  RELEASE=siglake &&
  cd "$CHECKOUT_ROOT/deploy/terraform/aws" &&
  terraform init &&
  terraform apply \
    -var "name=$NAME" \
    -var "region=$APPLY_REGION" &&
  REGION="$(terraform output -raw region)" &&
  terraform_output_is_set region "$REGION" &&
  CLUSTER="$(terraform output -raw cluster_name)" &&
  terraform_output_is_set cluster_name "$CLUSTER" &&
  SECRET_ARN="$(terraform output -raw rds_secret_arn)" &&
  terraform_output_is_set rds_secret_arn "$SECRET_ARN" &&
  EFS_ID="$(terraform output -raw efs_file_system_id)" &&
  terraform_output_is_set efs_file_system_id "$EFS_ID" &&
  TERRAFORM_VALUES_FILE="$(mktemp -t siglake-helm-values.XXXXXX.yaml)" &&
  terraform output -raw helm_values > "$TERRAFORM_VALUES_FILE" &&
  terraform_values_are_ready "$TERRAFORM_VALUES_FILE" &&
  cd "$CHECKOUT_ROOT" &&
  VALUES_FILE="$TERRAFORM_VALUES_FILE"
}

if terraform_bootstrap; then
  TERRAFORM_VALUES_FILE=
else
  BOOTSTRAP_STATUS=$?
  if [ -n "$TERRAFORM_VALUES_FILE" ]; then
    rm -f "$TERRAFORM_VALUES_FILE" || :
  fi
  REGION=
  CLUSTER=
  SECRET_ARN=
  EFS_ID=
  VALUES_FILE=
  if [ -n "$CHECKOUT_ROOT" ]; then
    cd "$CHECKOUT_ROOT" 2>/dev/null || :
  fi
  (exit "$BOOTSTRAP_STATUS")
fi

The full plan creates roughly 60 resources and takes about 15 minutes, most of it the EKS control plane.

Post-apply

  1. Write the cluster context to a fresh kubeconfig, confirm that kubectl reaches that cluster, then create the release namespace. The mktemp file starts empty, so no context already in your kubeconfig can answer for the new cluster. Each command runs only if the one before it succeeded: if the chain stops, nothing is created and $? is nonzero.
KUBECONTEXT="$CLUSTER"
KUBECONFIG="$(mktemp -t siglake-kubeconfig.XXXXXX)" &&
  export KUBECONFIG &&
  printf 'apiVersion: v1\nkind: Config\n' > "$KUBECONFIG" &&
  aws eks update-kubeconfig \
    --region "$REGION" \
    --name "$CLUSTER" \
    --alias "$KUBECONTEXT" \
    --kubeconfig "$KUBECONFIG" &&
  SELECTED_CONTEXT="$(kubectl config current-context)" &&
  if [ "$SELECTED_CONTEXT" != "$KUBECONTEXT" ]; then
    echo "kubeconfig selects $SELECTED_CONTEXT, not $KUBECONTEXT" >&2
    false
  fi &&
  kubectl --context "$KUBECONTEXT" get nodes &&
  NAMESPACE_MANIFEST="$(kubectl --context "$KUBECONTEXT" create namespace \
    "$NAMESPACE" --dry-run=client -o yaml)" &&
  printf '%s\n' "$NAMESPACE_MANIFEST" \
    | kubectl --context "$KUBECONTEXT" apply -f -

$KUBECONTEXT and $KUBECONFIG are what the rest of the procedure targets. Steps 3 to 7 name that context on every kubectl and helm call. The smoke test calls kubectl without a context argument, so it follows the current context in the exported KUBECONFIG, which is the context this step checked.

  1. Create the EFS StorageClass. The Terraform module installs the EFS CSI driver, but it does not create Kubernetes objects.
cat <<YAML | kubectl --context "$KUBECONTEXT" apply -f -
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: efs-sc
provisioner: efs.csi.aws.com
parameters:
  provisioningMode: efs-ap
  fileSystemId: $EFS_ID
  directoryPerms: "700"
reclaimPolicy: Delete
volumeBindingMode: Immediate
YAML
  1. Copy the RDS credentials from AWS Secrets Manager into the namespace. The Secret name matches postgres.existingSecret in the generated Helm values. Write the converter to a temporary file first, because the secret JSON takes standard input. The JSON stays off every command line, and the converter prints a manifest only after it finds usable values for all five fields. A failed retrieval or malformed payload sends nothing to kubectl.
(
SECRET_SCRIPT=""
SECRET_JSON=""
SECRET_MANIFEST=""

cleanup_secret() {
  SECRET_STATUS=$?
  trap - 0 HUP INT TERM
  unset SECRET_JSON SECRET_MANIFEST
  if [ -n "$SECRET_SCRIPT" ] && ! rm -f "$SECRET_SCRIPT"; then
    [ "$SECRET_STATUS" -ne 0 ] || SECRET_STATUS=1
  fi
  exit "$SECRET_STATUS"
}
trap cleanup_secret 0
trap 'exit 1' HUP INT TERM

SECRET_SCRIPT="$(mktemp -t siglake-secret.XXXXXX.py)" || exit
cat > "$SECRET_SCRIPT" <<'PY' || exit
import base64
import json
import sys

REQUIRED = ("host", "port", "user", "password", "database")

namespace, name = sys.argv[1:3]
try:
    payload = json.loads(sys.stdin.read())
except ValueError:
    sys.exit("secret payload is not valid JSON")
if not isinstance(payload, dict):
    sys.exit("secret payload is not a JSON object")
missing = [
    key
    for key in REQUIRED
    if key not in payload
    or (isinstance(payload[key], str) and not payload[key].strip())
]
if missing:
    sys.exit(f"secret payload is missing: {', '.join(missing)}")

unusable = []
for key in REQUIRED:
    value = payload[key]
    if key == "port":
        valid_type = isinstance(value, str) or (
            isinstance(value, int) and not isinstance(value, bool)
        )
    else:
        valid_type = isinstance(value, str)
    if not valid_type:
        unusable.append(key)
if unusable:
    sys.exit(f"secret payload has unusable fields: {', '.join(unusable)}")

lines = [
    "apiVersion: v1",
    "kind: Secret",
    "metadata:",
    f"  name: {name}",
    f"  namespace: {namespace}",
    "type: Opaque",
    "data:",
]
for key in REQUIRED:
    encoded = base64.b64encode(str(payload[key]).encode()).decode()
    lines.append(f"  {key}: {encoded}")
print("\n".join(lines))
PY

if SECRET_JSON="$(aws secretsmanager get-secret-value \
  --region "$REGION" \
  --secret-id "$SECRET_ARN" \
  --query SecretString \
  --output text)"; then
  :
else
  SECRET_STATUS=$?
  echo "Secrets Manager retrieval failed; no Secret applied." >&2
  exit "$SECRET_STATUS"
fi

SECRET_MANIFEST="$(
  printf '%s' "$SECRET_JSON" \
    | python3 "$SECRET_SCRIPT" "$NAMESPACE" "${NAME}-postgres"
)" || exit
printf '%s\n' "$SECRET_MANIFEST" \
  | kubectl --context "$KUBECONTEXT" apply -f -
)
  1. Confirm that the StorageClass and Secret exist before installing the chart.
kubectl --context "$KUBECONTEXT" get storageclass efs-sc
kubectl --context "$KUBECONTEXT" --namespace "$NAMESPACE" \
  get secret "${NAME}-postgres"
  1. Install the chart from the checkout root. The generated values configure Postgres, S3 and IRSA. The explicit override selects the EFS StorageClass for the write-ahead log (WAL). If Helm fails, the block exits with Helm's status and keeps the values file, so you can fix the cause and run this step again. Stop there: do not continue to verification. The values file is deleted only after Helm reports success.
(
if helm upgrade --install "$RELEASE" "$CHECKOUT_ROOT/deploy/helm/siglake" \
  --kube-context "$KUBECONTEXT" \
  --namespace "$NAMESPACE" \
  --values "$VALUES_FILE" \
  --set wal.storageClassName=efs-sc \
  --wait --timeout 10m; then
  rm -f "$VALUES_FILE"
else
  HELM_STATUS=$?
  printf 'helm install failed; values kept at %s\n' "$VALUES_FILE" >&2
  exit "$HELM_STATUS"
fi
)
  1. Verify the deployment. Wait for the ingester Deployment and query StatefulSet, then run the smoke test from the checkout root.
(
  set -e
  kubectl --context "$KUBECONTEXT" --namespace "$NAMESPACE" rollout status \
    "deployment/${RELEASE}-ingester" --timeout=5m
  kubectl --context "$KUBECONTEXT" --namespace "$NAMESPACE" rollout status \
    "statefulset/${RELEASE}-query" --timeout=5m
  cd "$CHECKOUT_ROOT"
  env "SIGLAKE_"RELEASE="$RELEASE" "SIGLAKE_"NAMESPACE="$NAMESPACE" \
    KUBECONFIG="$KUBECONFIG" ./deploy/aws/smoke.sh
)

The outputs you wire into chart values:

Output Feeds
serviceaccount_role_arn serviceAccount.annotations."eks.amazonaws.com/role-arn"
warehouse_bucket s3.bucket
region s3.region
efs_file_system_id The post-apply StorageClass.

helm_values sets postgres.existingSecret to <name>-postgres. The command above creates that Secret from the module's rds_secret_arn output before Helm starts the pods.

Helper scripts

deploy/aws/ wraps the common flows:

Use up.sh instead of the numbered procedure. Running both repeats the Helm installation and the Kubernetes setup.

Script Purpose
up.sh Automated alternative to this procedure: apply Terraform, configure the cluster, create Secrets and StorageClasses, install Helm and wait for readiness.
down.sh Remove the release and application resources. Keep the EKS cluster and its supporting VPC, EFS and ECR resources by default.
smoke.sh End-to-end correctness check.
query-bench.sh Query benchmark suite.
warehouse-load.sh Load a corpus into the warehouse.
operator-smoke.sh Operator path validation.

deploy/aws/config/ holds ready-made values files: values.smoke.yaml, values.multi-replica.yaml and values.query-perf.yaml.

Adjust the defaults for production

The defaults are sized for validation runs. Change these before you carry production traffic.

Raise rds_instance_class. The catalog is on the commit path for every drain batch, and commit duration shows up directly in siglake_iceberg_commit_duration_seconds. Size it against your commit rate.

Add a NAT gateway per availability zone if you need tolerance for the loss of one zone. The module provisions a single NAT gateway, which is a single-zone dependency.

Move EFS to provisioned or elastic throughput for heavy workloads. The module provisions bursting throughput; under sustained high ingest, burst credits deplete and WAL writes slow down.

Decide on the S3 lifecycle rule deliberately. Glacier-IR transitions are available through warehouse_lifecycle_days_to_glacier and off by default. Transitioning warehouse objects raises query latency on cold data, because Siglake does not stage data back.

Request an EC2 vCPU quota increase ahead of time. Multi-ingester topologies hit vCPU service quotas before anything else. Several benchmark rounds failed with VcpuLimitExceeded even though the arithmetic fitted, because terminating instances from an earlier teardown still counted against the quota. Leave a cooldown after a teardown for the same reason.

Teardown

Run the cleanup script from the same shell, so it inherits the KUBECONFIG step 2 exported and removes the release from the cluster you installed it on. The block deletes the kubeconfig only after the script reports success.

(
  cd "$CHECKOUT_ROOT" || exit
  if ./deploy/aws/down.sh; then
    rm -f "$KUBECONFIG"
  else
    DOWN_STATUS=$?
    printf 'cleanup failed; kubeconfig kept at %s\n' "$KUBECONFIG" >&2
    exit "$DOWN_STATUS"
  fi
)

If the directory change or down.sh fails, the block exits with that status and keeps the kubeconfig. Fix the cause, then run the same block again from the same shell. The retry needs that kubeconfig to reach the cluster, and the script skips the Helm release, Secret, PVCs and namespace that the first run already deleted.

down.sh exits with Terraform's own status when the final terraform destroy fails, and prints down complete only after a destroy that succeeded. A failed destroy leaves AWS resources up and billing. The script also fails before it reaches Terraform, for a missing Terraform directory or an unknown cleanup mode, so read its last error line to tell the two cases apart. The Helm and kubectl steps ahead of the destroy are best effort on purpose: the script carries on when they fail, and their failures do not change its exit status.

The default cluster cleanup mode removes the Helm release, namespace, RDS instance, warehouse bucket and application IAM resources. It keeps the EKS cluster, VPC, EFS and ECR resources warm for the next run. The script empties the versioned warehouse bucket in this mode, including its object versions.

Warning

The script's all cleanup mode deletes the complete Terraform stack. If you also enable warehouse emptying, it deletes every object version first.