Data Architect - Austin

Biorce
Biorce

IT

Austin, TX, USA

Posted on Aug 12, 2026

About the company

Biorce is a pioneering Healthtech company dedicated to revolutionizing drug development through the power of AI. We are passionate about accelerating medical advancements and improving patient outcomes.

Our team comprises seasoned clinical research professionals, data scientists, and AI experts, working collaboratively to bridge the gap between cutting-edge technology and real-world clinical needs.

With an unwavering commitment to revolutionize healthcare, we envision a world where all patients benefit from accelerated and cost-effective access to treatments. Biorce is poised to redefine the landscape of healthcare, shaping a future where innovation and accessibility converge for the betterment of humanity.

About the role

Following our successful expansion into the U.S. and continued growth across Europe, we are seeking a Senior Data Engineer / Data Architect to lead a greenfield build-out of Biorce's agentic data platform from our Austin hub.

This is a rare opportunity to design a platform from a clean slate. You will architect and build a hybrid, open-source + Databricks Lakehouse platform on Google Cloud (GCP), with dbt as the transformation and modeling layer, deliberately favoring open standards and open table formats to stay flexible and avoid lock-in, while leaning on Databricks for scale where it earns its place.

At the center of the vision is a self-healing, agentic data warehouse: a platform where querying, pipeline construction, and the merging and mapping of heterogeneous sources into a common data model are all driven by AI agents, with humans setting direction, guardrails, and standards. You will be the technical anchor who turns that vision into a running system.

This is a high-impact, hands-on leadership role for someone who wants to design the data backbone of a next-generation clinical AI platform and rethink how data engineering itself gets done.

Who We're Looking For

A senior engineer or architect who is equally comfortable whiteboarding a Lakehouse from first principles and getting into the weeds of a Spark job or a dbt model. Someone who has strong opinions on data architecture, cares deeply about reliability and compliance in a regulated environment, and is genuinely excited about using AI agents to accelerate not replace rigorous engineering.

You will work closely with data scientists, AI engineers, MLOps, and DevOps to design and operationalize robust data flows that fuel advanced analytics, machine learning, and regulatory-grade insights while mentoring the team and raising the bar on how we build.

The Self-Healing, Agentic Data Warehouse

This is the defining concept of the role. We want to build a platform where the core data engineering loops are agentic and self-correcting:

Agentic querying, a natural-language and self-correcting query layer over a well-defined semantic model, where agents translate intent into validated SQL, explain results, and recover from errors without hand-holding.

Agentic pipeline building, agents that scaffold, build, test, and document ingestion and transformation pipelines (Spark / dbt) from specifications, with engineers reviewing and steering rather than hand-writing boilerplate.

Agentic mapping to a common data model, agents that profile, merge, and map heterogeneous clinical, research, and third-party sources into a canonical common data model, handling schema mapping, entity resolution, and normalization.

Self-healing operations, pipelines and tests that detect schema drift, data-quality failures, and anomalies, then diagnose and auto-remediate (or open a well-scoped fix for review), minimizing manual firefighting.

You will design the guardrails, evaluation criteria, and human-in-the-loop checkpoints that make this safe and trustworthy for clinical, regulated data.

Key Responsibilities

Platform & Architecture (Greenfield)

  • Architect and build, from a clean slate, a hybrid open-source + Databricks Lakehouse platform on GCP, with dbt as the transformation standard.
  • Make deliberate build vs. buy and open-source vs. managed decisions, selecting and integrating open table formats, orchestration, and processing frameworks that keep us flexible and cost-aware.
  • Design scalable, fault-tolerant batch and streaming pipelines using Spark/PySpark, Structured Streaming, and Databricks Workflows / declarative pipelines.
  • Establish dbt project structure, layering (medallion / bronze-silver-gold), testing, documentation, and CI/CD via the dbt-Databricks adapter.
  • Define a common/canonical data model for clinical and research data and the mapping strategy into it.
  • Set standards for data quality, lineage, and observability across all systems through rigorous validation, testing, and monitoring.

AI-First / Agentic Data Engineering

  • Design and build the self-healing, agentic data warehouse described above, agentic querying, pipeline building, source mapping, and self-healing operations.
  • Pioneer and scale AI-assisted data engineering across the team using Claude Code, coding agents, and agentic workflows to accelerate pipeline, dbt model, and test development.
  • Build reusable context, patterns, and internal tooling (including MCP-style integrations where useful) that make the whole team faster and more consistent.
  • Define guardrails, review practices, and evaluation criteria so agent-generated code and mappings are validated to the standard required for clinical and regulated data, speed with accountability.

Engineering Excellence & Compliance

  • Develop and optimize SQL, Python, and PySpark transformations for performance, cost, and maintainability.
  • Manage storage, partitioning, clustering, and lifecycle strategies for efficiency and cost control (FinOps mindset).
  • Ensure compliance with SOC 2, ISO 27001, HIPAA, GDPR, and clinical data governance standards in all data operations, with strong access control, lineage, and auditability (e.g., Unity Catalog).
  • Champion infrastructure-as-code (Terraform, Databricks Asset Bundles) and GitOps-based, modular, version-controlled pipeline development.
  • Mentor engineers, lead technical design reviews, and shape the long-term evolution of Biorce's data and AI architecture.

Must-Haves

  • 6+ years in Data Engineering, with recent experience at a senior, lead, or architect level.
  • Demonstrated experience designing and building data platforms from scratch (greenfield), including architecture and technology-selection decisions.
  • Deep hands-on expertise with the Databricks Lakehouse: Delta Lake / open table formats, Spark/PySpark, Databricks Workflows, and SQL warehousing.
  • Strong production experience with dbt (Core and/or Cloud): modeling, testing, macros, sources/exposures, and deployment.
  • Hands-on experience running data workloads on Google Cloud (GCP).
  • Hands-on experience using AI coding assistants or agents (e.g., Claude Code) in real engineering work and a point of view on how to do it responsibly.
  • Comfort composing open-source data tooling (e.g., Spark, Airflow/Dagster, open table formats, Great Expectations) into a coherent platform.
  • Expert-level SQL and strong Python for large-scale data transformation and automation.
  • Proven track record designing batch and streaming pipelines with scalable, fault-tolerant architectures.
  • Solid grounding in data modeling, schema design, and Lakehouse optimization, including mapping disparate sources into a common/canonical data model.
  • Experience with infrastructure-as-code and CI/CD for data (Terraform, Databricks Asset Bundles, or GitOps).
  • Bachelor's or Master's in Computer Science, Engineering, or a related quantitative field (or equivalent experience).

Nice-to-Haves

  • Experience building agentic workflows or internal AI tooling (MCP servers, LLM-in-the-loop pipelines, natural-language-to-SQL, prompt/context engineering).
  • Experience with entity resolution, schema mapping, and normalization at scale.
  • Experience with clinical, biomedical, or healthcare datasets and common data models such as OMOP CDM or FHIR.
  • Familiarity with MLflow, Mosaic AI, Vertex AI, or ML metadata / feature-store tooling.
  • Data quality and validation frameworks such as Great Expectations, dbt tests, or TFX Data Validation.
  • Streaming at scale (Kafka / Pub/Sub, Apache Beam) and knowledge of Spark internals.
  • Containerized workflows (Docker, Kubernetes) and a strong focus on reliability, observability, and continuous improvement.
  • Experience with data cataloging and governance (Unity Catalog, Data Catalog, or similar).

Why Join Us?

  • Build a next-generation clinical data platform from a clean slate and define what a self-healing, AI-first data warehouse actually looks like in a regulated environment.
  • A dynamic work environment with an international team, where collaboration and diversity thrive.
  • Work alongside top talent, united by a shared purpose and committed to making a real impact.
  • Comprehensive private health coverage to support your physical and mental well-being.
  • Hybrid work model offering flexibility to balance your professional and personal life.
  • Company events to celebrate achievements and enjoy time together.
  • A MacBook and a best-in-class AI toolchain to maximize your productivity.
  • Our office is pet-friendly, you'll likely be greeted by a few wagging tails upon arrival.

--

By submitting this application, I agree that my personal data will be collected, processed, and retained by the company solely for the purposes of managing and assessing my candidacy.