Work · Co-founded venture
Discoveria: a clinical-trials platform, upgraded 2 → 3 → 4 → 5 → 6
A pre-launch, patient-friendly clinical-trials platform. I took it from Grails 2.5.6 to 6.2.3 one major version at a time, and built an AI content pipeline behind a ~2M-page site built from 500K+ clinical trials.
- 2.5.6
- 3.3.18
- 4.1.4
- 5.3.6
- 6.2.3
Discoveria, Inc. is a patient-friendly clinical-trials discovery platform built on ClinicalTrials.gov’s AACT dataset: 500K+ studies. I co-founded it, and I lead all of its engineering.
The problem
ClinicalTrials.gov lists 500K+ studies, but the data is dense and technical, and it’s hard for patients and families to understand or search. Discoveria mirrors the full AACT research database and layers patient-friendly, AI-assisted content and search on top.
The original 2020 build ran on Grails 2.5.6 / Java 8. Before building any further, I brought the whole stack up to date.
What I built
- A stepwise upgrade through every major Grails version: 2.5.6 → 3.3.18 → 4.1.4 → 5.3.6 → 6.2.3 (Java 8 → 11 → 17) in about six months, Sep 2024 → Mar 2025. Each step had its own theme:
- 2 → 3: the big rewrite (Gradle, Spring Boot, YAML config, the new plugin system).
- 3 → 4: Java 11 and GORM 7.
- 4 → 5: Gradle 7, Groovy 3, Micronaut.
- 5 → 6: Java 17 and Spring Security 6.
- Multi-datasource GORM: the app database, a mirror of the 52-table AACT schema, and an upstream source read through raw SQL. Transactions are tied to a specific datasource, and lazy-loading traps are designed out.
- An AI content engine that rewrites dense trial data and generates patient-friendly content and SEO metadata (JSON-LD / Open Graph). It’s the AI content pipeline behind a ~2M-page site built from 500K+ clinical trials:
- an OpenAI prompt library, with a dedicated rate limiter (3,500 RPM / 200K TPM, covered by tests) and retries with backoff
- an OpenAI Batch API pipeline that submits, polls and imports results
- validation of AI-generated metadata before it’s used
- content-sensitivity levels (general, medical professional, patient-specific, sensitive, pediatric, mental health) with human review queues
- per-entity processing for conditions, drugs and interventions, treatments and studies, plus a nightly SEO job and stale-content detection
- External API integrations: OpenAI, the ClinicalTrials.gov (AACT) data, a trial-sponsor API, Stripe webhooks, MailChimp and Bitly. All run through a retry helper with API performance tracking, alongside bot detection and REST API authentication.
- The data-sync pipeline: re-engineered with batched flushing, parallel chunked processing, explicit executors and JVM tuning. Benchmarked at up to 83× the original throughput.
- Infrastructure as code: Terraform. Deploys go GitHub Actions → S3 → SSM with a health-checked rollback.
Results
- Every major Grails version, one at a time, Java 8 → 17, in about six months.
- Scale: 111 domain classes and 103 services, about 175–180K lines of first-party Groovy/GSP, over a 500K+ study dataset.
- AI at scale, done responsibly: rate limits, the Batch API, output validation, sensitivity levels and human review, not a toy chatbot wrapper.
- Data sync benchmarked at up to 83× the original throughput.
- Built for AI-assisted engineering: a CLAUDE.md with a ~2K-line reference library, multi-session playbooks, and a two-tier TODO / session-continuity workflow.
My role
Co-founder. I own the architecture, the framework upgrade, the data model and sync pipeline, the AI content system, and the infrastructure and deployment.