Origin AI

Research Data Platform

2021-02 → 2025-10 · Software Engineer → Lead → Senior

Internal cloud systems for a research organisation: secure field data storage, a cloud-native analytics platform, and the batch processing layer behind the AI servers.

Composition

  • ComputeAWS Lambda · ECS · ECR · EC2
  • OrchestrationAWS Step Functions · Glue
  • StorageAurora · Redshift · DynamoDB · S3
  • InterfaceReact · Next.js · Redux · TypeScript
  • AnalysisMATLAB integration · Athena
  • DeliveryJenkins · Jest · Git hooks
  • TransportMQTT · SQS/SNS · Kafka

Mission & Constraints

  1. The starting condition was a research team losing seven days to every data cycle. The design began with user research rather than architecture — what researchers actually did with field data, and where the days were going.

  2. The resulting system became the standard for field data across research teams, which is the part that mattered: an internal tool nobody adopts has no measurement to report.

  3. Later work moved down the stack — Jenkins pipelines and dual environments, Jest and CI groundwork, Git hooks across departments — and then into the AI servers themselves, where batch processing took out half the CPU and most of the cost.

Implementation

Compute

AWS Lambda · ECS · ECR · EC2

Orchestration

AWS Step Functions · Glue

Storage

Aurora · Redshift · DynamoDB · S3

Interface

React · Next.js · Redux · TypeScript

Analysis

MATLAB integration · Athena

Delivery

Jenkins · Jest · Git hooks

Transport

MQTT · SQS/SNS · Kafka

Measured Outcomes

Research data processing time

−95.2 %

7 days to8h

Origin AI — internal research tooling

Field data parsing time

−95 %

baseline toresidual

Cloud-native analytics platform, AWS + MATLAB

AI server compute cost

−63 %

baseline toresidual

Origin AI — batch processing on AI servers

CPU utilisation

−52 %

baseline toresidual

Origin AI — batch processing on AI servers

Inference speed

+41 %

baseline toachieved

Origin AI — batch processing on AI servers

Questions about this work?

Happy to go deeper on the architecture, the constraints, or how the numbers were measured.