About SandboxAQ
SandboxAQ is a high-growth company delivering AI solutions that address some of the world's greatest challenges. The company’s Large Quantitative Models (LQMs) power advances in life sciences, financial services, navigation, cybersecurity, and other sectors.
We are a global team that is tech-focused and includes experts in AI, chemistry, cybersecurity, physics, mathematics, medicine, engineering, and other specialties. The company emerged from Alphabet Inc. as an independent, growth capital-backed company in 2022, funded by leading investors and supported by a braintrust of industry leaders.
At SandboxAQ, we’ve cultivated an environment that encourages creativity, collaboration, and impact. By investing deeply in our people, we’re building a thriving, global workforce poised to tackle the world's epic challenges. Join us to advance your career in pursuit of an inspiring mission, in a community of like-minded people who value entrepreneurialism, ownership, and transformative impact.
The Opportunity
The AI Sim R&D team builds leading-edge ML and physics-based models ("LQMs") to advance drug discovery. Within this team, AQCell is our virtual cell platform: it takes a cell representation (e.g. basal gene expression) and a perturbation descriptor (e.g. SMILES, dose, and time), predicts the resulting transcriptomic response, and maps that response through pathway activity to functional endpoints such as cell viability, IC50, and toxicity dose-response — helping drug discovery scientists understand the biological repercussions of a compound across the cell, not just whether it binds its target.
As a Machine Learning Engineer on AQCell, you will build and maintain the models and data infrastructure that power this pipeline. You will work across large, heterogeneous transcriptomic and functional-endpoint datasets (e.g. LINCS L1000, GDSC, Tahoe-100M, and DILImap), train and evaluate expression-perturbation and cell-viability prediction models over them, and help harden our evaluation pipeline and baselines so we can trust and improve model performance over time. This role sits at the intersection of machine learning and biology, and offers the opportunity to shape how virtual cell models are trained, validated, and scaled toward real drug discovery decisions.
Key Responsibilities
Model Development: Build, train, and maintain machine learning models for expression-perturbation prediction (e.g. transcriptomic response to a drug perturbation) and downstream functional-endpoint prediction (e.g. cell viability, IC50, toxicity dose-response).
Large-Scale Dataset Management: Acquire, harmonize, and manage large-scale biological datasets (e.g. LINCS L1000, GDSC, Tahoe-100M, DILImap) — including schema harmonization, normalization, and de-duplication across cell, drug, and assay identifiers — and manage model training pipelines over these pooled datasets.
Evaluation & Baselines: Contribute to automating and hardening the end-to-end evaluation pipeline, including implementing robust statistical baselines (e.g. cell- and drug-conditioned mean baselines) to rigorously benchmark model performance.
Research Translation: Translate ideas from the scientific literature (e.g. transformer-based perturbation models, knowledge graph and GNN-based embeddings) into working, well-tested code integrated into our modeling framework.
Cross-Functional Collaboration: Partner with computational biologists, software engineers, and product stakeholders to ensure models are grounded in sound biology and are usable in real drug discovery workflows.
Communication: Clearly document methods, assumptions, and results, and communicate findings to both technical and non-technical stakeholders.
Essential Skills & Experience
Academic Foundation: Bachelor's degree in a scientific or quantitative field (Computer Science, Physics, Mathematics, Biology, Chemistry, or related); an advanced degree (MS or PhD) is preferred.
Applied ML in Science: Demonstrated experience building and maintaining machine learning models in a scientific discipline in an industry setting, including taking models from prototype through validation and maintenance; experience in a life sciences setting is preferred.
Large-Scale Data Management: Required experience managing large-scale datasets and managing the training of models over them, including data ingestion, cleaning, versioning, and pipeline maintenance at scale.
Software Engineering: Strong Python programming skills and experience with modern ML frameworks (e.g. PyTorch, JAX) and experiment tracking/data versioning tools (e.g. Weights & Biases).
Scientific Rigor: Ability to design sound evaluation methodology (e.g. train/test splitting strategies, held-out generalization tests) and to critically interpret model performance against meaningful baselines.
Highly Desired Skills & Experience
Bioinformatics & Computational Biology: Experience with bioinformatics and computational biology data analysis tools, particularly transcriptomics harmonization and normalization tools (e.g. batch correction, pseudobulking, gene ID mapping/standardization).
Domain Datasets: Familiarity with public perturbation or drug-sensitivity datasets such as LINCS L1000, GDSC, DepMap, or single-cell perturbation atlases.
Knowledge Graphs & GNNs: Experience with knowledge graph embeddings or graph neural networks applied to drugs, targets, or cells.
Cheminformatics: Familiarity with cheminformatics representations (SMILES, InChI keys, PubChem/Cellosaurus identifiers) used to harmonize drug and cell metadata across datasets.
Collaborative Innovation: Experience working in interdisciplinary environments where AI intersects with the biological sciences.
Why Join Us?
We offer competitive compensation, a comprehensive benefits package, and opportunities for professional growth.
Compensation: Competitive base salary, performance-based incentives or bonuses (where applicable), and equity participation.
Benefits: Comprehensive medical, dental, and vision coverage for employees and dependents with generous employer premium contributions, retirement savings with company matching, paid parental leave, and inclusive family-building benefits.
Work-Life Balance: Flexible paid time off, company-wide seasonal breaks, and support for flexible work arrangements that enable sustainable performance.
Career Development: Opportunities for continuous learning and growth through on-the-job development, cross-functional collaboration, and access to internal learning and development programs.
SandboxAQ Welcomes All
We are committed to fostering a culture of belonging and respect, where diverse perspectives are actively sought and valued. Our multidisciplinary environment provides ample opportunity for continuous growth - working alongside humble, empowered, and ambitious colleagues ready to tackle epic challenges.
Equal Employment Opportunity: All qualified applicants will receive consideration regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, or Veteran status.
Accommodations: We provide reasonable accommodations for individuals with disabilities in job application procedures for open roles. If you need such an accommodation, please let a member of our Recruiting team know.
Read: Guidance for candidates on using AI Tools in interviews
