Cipher pulls encrypted bulk files, decrypts them in memory, and streams the rows into a warehouse. The heavy files route to Fargate automatically — so a fifteen-minute function ceiling never stops a multi-gigabyte job. Run a sample below, or upload your own encrypted file and watch it decrypt in your browser.
Pick a sample to watch the pipeline run at scale (simulated). Your-file mode really decrypts a small file in your browser — nothing is uploaded.
It's the big-file path — here's why it isn't a bigger Lambda, and why it isn't Glue.
Past a certain size, an encrypted file hits Lambda's hard 15-minute ceiling long before it runs low on memory. Give the function the maximum RAM and the CPU that comes with it — it still won't matter. A streaming decrypt-and-load holds only a few rows at a time, so peak memory stays low (a few hundred MB); what runs out is wall-clock time, and that ceiling can't be raised.
And nothing bottlenecks on the output side: rows are written as Parquet straight to S3 and queried in place by Athena — there's no ingest channel or commit step to throttle. The one hard limit left is the function's wall-clock budget.
The only lever left is removing the time limit. Fargate runs the exact same container with no ceiling — and a completed run costs less than a Lambda that times out at 15 minutes, retries, and delivers zero rows.
if a job is too big for Lambda, the reflex is AWS Glue — here's why it's the wrong fit, four reasons
gpg system binary, and Glue Spark would force a decrypt-to-plaintext-at-rest step — breaking the security model.Ciphertext is piped through gpg in a memory stream straight into the parser. The cleartext is never staged to S3 or written to the container filesystem — it exists only in flight. The upload demo does the same thing in your browser.
A single container image is deployed as both a Lambda (container package) and a Fargate task. No second codebase to drift — the small-file and big-file paths are byte-for-byte identical logic.
The orchestrator keeps a manifest of path + size + last-modified. Unchanged files are skipped before anything downloads. Re-run the same file in simulated mode — it stops at the orchestrator instead of reloading.
Columns load as unbounded strings, so nothing fails on a type mismatch. A data-quality scan samples the rows and sends one summary per file flagging mis-typed columns — without ever dropping the load.