customers, contracts, invoices, payments, train-schedules, locomotives,
trains and wagons. 319 fields across the nine datasets, all reusing the
existing engine — no change to export.types.ts was needed, which is the
result the bookings-first phase was meant to test.
Per-dataset notes worth keeping:
- trains resolves route, stations and current yard, which the list endpoint
never loads — the UI shows raw FK uuids there today.
- wagons reads tare/payload/length off wagon_types (they are not on the
wagon), and reproduces the service's attachStatusDates() as correlated
subqueries. wagon_status_logs stores from_status/to_status, not status.
- payments applies no soft-delete guard: freight.payments has neither
deleted_at nor updated_at, so the usual predicate is a 42703. Failure
columns are failer_code/failer_message. payment_refunds stores MINOR
units, so refundedTotal divides by 100.
- train-schedules derives freightType from the bookings aboard rather than
a column, matching the list service.
- customers stays one row per company; profiles, bookings and invoice
totals aggregate in subqueries. Verified no row multiplication: trains,
customers and contracts each return exactly their counted row count while
selecting one-to-many aggregate fields.
EXPLAIN-validated against the database: every dataset's widest query, its
count query, and all 319 fields individually. That run caught five columns
typed varchar rather than timestamp (companies.date_registered,
renewal_date, renewed_from, renewed_to and invoices.eims_ack_date), which
were being pushed through to_char and would have 500'd the moment anyone
ticked them; they now export verbatim.
All nine count endpoints verified equal to SELECT count(*) on their table.
Adds a parallel export system the reports module can also draw on. A dataset
describes a table's exportable fields — including related-entity detail the
list page never shows — and the engine assembles a query from whichever fields
the caller picked.
GET /exports catalog (metadata only; select/requires never ship)
GET /exports/:key/count exact row count + per-format caps
GET /exports/:key/download csv | xlsx | pdf
Two invariants carry the design:
- Every lazy join is a LEFT join, and ExportJoin has no 'kind' field to make
anything else expressible. An inner join added because a checkbox was ticked
would change the rowset, so two exports of the same filters would disagree on
their row count.
- Because of that, the count cannot depend on field selection, so /count runs
base + alwaysJoin only and is exact rather than an estimate. Verified: count
and the delivered file both report 223 rows.
One-to-many relations (a booking's containers) aggregate in a correlated
subquery rather than joining, so a row can never multiply.
Export rides each dataset's existing view permission — no new permission keys
and no seeder change. Sensitive columns are simply never declared as fields:
raw gateway payloads, signature blobs, error dumps, raw jsonb snapshots,
internal user UUIDs and review notes are all absent by construction.
bookings ships 77 fields across 10 groups. scripts/validate-export-datasets.ts
EXPLAINs every dataset's widest query, its count query, and each field on its
own against the real database — the per-field pass is what catches a field
referencing a join it forgot to declare, which otherwise only fails when that
one field is picked alone.