Secure data ingestion · AWS-native

Encrypted in.
Clean rows out.

Cipher pulls encrypted bulk files, decrypts them in memory, and streams the rows into a warehouse. The heavy files route to Fargate automatically — so a fifteen-minute function ceiling never stops a multi-gigabyte job. Run a sample below, or upload your own encrypted file and watch it decrypt in your browser.

25M rows
largest file
2.1GB → ~20GB
encrypted → decrypted
0bytes
plaintext to disk
2runtimes
one codebase
01

Run the pipeline

Pick a sample to watch the pipeline run at scale (simulated). Your-file mode really decrypts a small file in your browser — nothing is uploaded.

cipher · ingest sim · no backend
1 · pick an encrypted file
2 · or drag the encrypted size
18 MB
size gate · runtime decision
routes to
Lambda
Under the 80 MB threshold — finishes inside the 15-minute function limit.
drop a file to ingest for real
⤓
Drop a .csv or encrypted .gpg / .ge
A plaintext .csv gets encrypted to the demo key first, then decrypted — so you see the whole lifecycle. An encrypted file is decrypted directly.
↓ Download demo public key — encrypt your own file to it
🔒 Your file never leaves your browser. Encryption and decryption run locally via OpenPGP.js — the same in-memory-only property the real pipeline relies on.
size gate · runtime decision
your file routes to
Lambda
Most files you'll drop are small, so they ingest inline on Lambda. Files over 80 MB would route to Fargate.
🌐
source
Web upload
⬆
https · presigned
Encrypted upload
λ
orchestrator
Manifest check
🧬
s3 · encrypted
Landing
📤
sqs
Queue
◇
size gate
Route by size
⚙
ingest · lambda
Decrypt + stream
▦
data warehouse
Committed
event log
warehouse · INGEST.table
0 rows
Encrypted payload — locked.
Pick a file and press Ingest.
02

Why Fargate?

It's the big-file path — here's why it isn't a bigger Lambda, and why it isn't Glue.

It's not memory-bound. It's time-bound.

Past a certain size, an encrypted file hits Lambda's hard 15-minute ceiling long before it runs low on memory. Give the function the maximum RAM and the CPU that comes with it — it still won't matter. A streaming decrypt-and-load holds only a few rows at a time, so peak memory stays low (a few hundred MB); what runs out is wall-clock time, and that ceiling can't be raised.

And nothing bottlenecks on the output side: rows are written as Parquet straight to S3 and queried in place by Athena — there's no ingest channel or commit step to throttle. The one hard limit left is the function's wall-clock budget.

The only lever left is removing the time limit. Fargate runs the exact same container with no ceiling — and a completed run costs less than a Lambda that times out at 15 minutes, retries, and delivers zero rows.

ingest time vs file size · the 15-min wall
ingest time → 18 MB 80 MB 425 MB → Fargate 2.1 GB → Fargate 15-min Lambda ceiling
Lambda (under wall)Fargate (over wall)

And why not Glue?

if a job is too big for Lambda, the reflex is AWS Glue — here's why it's the wrong fit, four reasons

  • One image, two runtimes. The Fargate task runs the exact container the Lambda path runs — a single codebase. Glue would be a third, separate runtime to write and keep in sync.
  • In-memory decryption. Files are GPG-decrypted in memory and never written as plaintext. Glue Python Shell can’t run the gpg system binary, and Glue Spark would force a decrypt-to-plaintext-at-rest step — breaking the security model.
  • Right shape for the work. This is single-stream, row-by-row decrypt → Parquet, not distributed batch over partitioned data. Spark’s parallelism buys nothing here and adds cluster overhead.
  • Cost. For this workload Fargate runs at roughly a third to half the price of Glue — you pay per-second for one right-sized task, not for a Spark cluster.
03

Other decisions behind it

01 decrypt path

Plaintext never touches disk

Ciphertext is piped through gpg in a memory stream straight into the parser. The cleartext is never staged to S3 or written to the container filesystem — it exists only in flight. The upload demo does the same thing in your browser.

staged plaintext objects: 0
02 one image, two runtimes

The same code runs both paths

A single container image is deployed as both a Lambda (container package) and a Fargate task. No second codebase to drift — the small-file and big-file paths are byte-for-byte identical logic.

images to maintain: 1
03 change detection

A manifest, not a re-pull

The orchestrator keeps a manifest of path + size + last-modified. Unchanged files are skipped before anything downloads. Re-run the same file in simulated mode — it stops at the orchestrator instead of reloading.

dedup key: path · size · mtime
04 load safety

One bad cell can't kill the file

Columns load as unbounded strings, so nothing fails on a type mismatch. A data-quality scan samples the rows and sends one summary per file flagging mis-typed columns — without ever dropping the load.

batched at 10K rows · Parquet row groups
04

What's underneath

S3 encrypted landing + curated SQS + DLQ Lambda size-gate + light ingest ECS Fargate heavy ingest Secrets Manager PGP key Parquet + Athena warehouse Glue catalog Auto-cleanup self-erasing loads (T+5m) SNS alerts + data quality