Build the Future of
Enterprise AI Data
Join the team rewriting the rules of post-training AI infrastructure. At Alchedata, you won't execute a playbook — you'll help write it.
Every Failure. A Better Checkpoint.
Alchedata builds the self-improving data layer for Physical AI. We help robotics, VLM, world model, and VLA teams close the gap between real-world failure and the next model improvement by turning every rollout into structured evidence.
Our platform connects failure diagnosis, targeted evaluation suites, evolving data recipes, agentic data operations, replay evidence, and validated simulation environments inside one governed loop: observe the failure, update the data plan, test the fix, and promote only when the checkpoint improves without breaking regression gates.
From multimodal collection and curation to portable simulation packages and checkpoint validation, Alchedata gives Physical AI teams the infrastructure to learn continuously from the long tail instead of treating each failure as a one-off debugging task.
AlcheEval
Turns long-tail failures into targeted eval suites, regression signals, and checkpoint scores.
AlcheRecipe
Continuously refines data mix, task balance, synthetic-real ratios, and curriculum sequencing.
AlcheCopilot
Runs agentic data operations for collection, multimodal annotation, curation, and anomaly review.
AlcheSims
Builds reproducible simulation environments from text, images, video, and verified assets.
Where You'll Make an Impact
Six high-impact roles with outsized scope, ownership, and direct product impact.
About the Role
We are hiring a founding Business Development & Sales Partner who will work directly with the CEO and CTO to drive our US market expansion. This is not a quota-carrying individual contributor role in a large sales org — this is a seat at the table. You will be one of the first commercial hires, helping define how Alchedata goes to market, acquires enterprise customers, and builds strategic partnerships.
This is a dual-function role: you will close deals as a direct seller and open doors through channel partnerships and ecosystem development. You will own the full deal lifecycle — from identifying high-value prospects to signing Statements of Work — while simultaneously building the partner infrastructure that multiplies our reach.
You are the right fit if you thrive in ambiguity, love translating technical complexity into business value, and want to be part of building something from the ground up. This role offers disproportionate impact, equity, and career trajectory for the right person.
Direct Sales Responsibilities
- Identify, qualify, and prioritize high-value enterprise prospects — specifically AI teams building LLMs, vertical agents, or specialized models in robotics and healthcare.
- Lead in-depth technical discovery sessions covering RLHF/RLAIF workflows, evaluator stacks, and model alignment pipelines to uncover genuine business need.
- Own the full sales cycle from initial demo through proposal, negotiation, and Statement of Work (SOW) signature, structuring deals for long-term expansion.
- Maintain clean CRM hygiene and deliver accurate pipeline forecasts to leadership.
Business Development Responsibilities
- Identify and engage strategic partners including AI consultancies, Systems Integrators (SIs), and complementary data platform providers.
- Experiment with reseller models and develop co-marketing initiatives with tooling providers (e.g., vector databases, MLOps platforms).
- Bring structured market intelligence back to the product and engineering teams to shape the roadmap for agentic automation.
Success Metrics
- Build a qualified pipeline of enterprise opportunities.
- Complete discovery calls with target ICP accounts.
- Establish active partnership conversations.
Qualifications
- 3+ years of full-cycle B2B sales experience selling technical products to enterprise buyers, with a track record of closing.
- Technical fluency in AI/ML concepts — you can credibly discuss the difference between pre-training and post-training, and hold your own in a conversation with an ML engineer.
- Excellent consultative communication skills with the ability to build trust, navigate complex stakeholder environments, and drive alignment across organizations.
- Startup DNA: self-directed, comfortable with ambiguity, and energized by building from scratch rather than executing a playbook.
- Willingness to travel for customer sites, conferences, and partner meetings.
Nice to Have
About the Role
Alchedata is building Data Infra 2.0 - the agent-orchestrated data platform purpose-built for next-generation AI systems. We integrate evaluation, post-training, and domain-specific RL environments into one continuous intelligence loop so frontier models can move from research demos to reliable real-world deployment.
We are looking for a highly motivated RL Environment Engineer Intern (VLM) to help design, implement, and optimize reinforcement learning environments for Vision-Language Models (VLMs) and multimodal agents. You will work on the environments that shape how models perceive, reason, act, and improve through feedback.
This is a hands-on role with real ownership. You will build high-quality, multi-modal environments used for post-training, evaluation, and data flywheel generation across our platform.
What You'll Do
- Design and implement domain-specific RL environments for VLM and multimodal agent tasks, such as visual reasoning, grounding, tool use, browser/computer interaction, document understanding, and agentic workflows.
- Build and extend multi-modal observation and action spaces spanning images, text, structured state, interface signals, and model/tool outputs.
- Design and implement reward models and reward functions for VLM and multimodal agent environments, including preference-aligned scoring, trajectory evaluation, termination conditions, validation logic, and curriculum learning pipelines.
- Integrate RL environments with our post-training stack, including SFT, preference optimization, RLHF, PPO-style training loops, and evaluation harnesses.
- Run large-scale experiments to benchmark environment quality, task difficulty, reward reliability, and generalization across model families.
- Partner closely with ML engineers and the founding team to turn research ideas into production-grade infrastructure, code, and documentation.
Requirements
- Currently pursuing a Master's or PhD degree in Computer Science, AI/ML, Robotics, Applied Math, or a related field, graduating in 2026 or later.
- Strong proficiency in Python; familiarity with PyTorch and modern ML tooling.
- Hands-on experience with at least one RL framework or environment stack such as Gymnasium, Stable-Baselines3, RLlib, CleanRL, or equivalent.
- Solid understanding of RL fundamentals, including MDPs, reward design, policy optimization, value functions, exploration, and credit assignment.
- Experience building evaluation pipelines, simulation environments, or agent systems for ML applications.
- Strong engineering instincts, including writing clean code, debugging experiments, and iterating quickly.
Nice to Have
What You'll Gain
- Ownership of production RL environments that directly shape how multimodal models are trained and evaluated.
- Exposure to the full stack: data generation, post-training, evaluation, and agent orchestration.
- The chance to work closely with a fast-moving founding team on cutting-edge VLM and agent infrastructure.
- A high-impact internship with the potential for a return offer, strong recommendation, and meaningful authorship on shipped systems.
About the Role
Alchedata is building Data Infra 2.0 — the agent-orchestrated data platform purpose-built for Physical AI. We integrate evaluation, post-training, and domain-specific RL environments into one continuous intelligence loop so the next generation of embodied AI and vision-language-action models can move from lab demos to real-world deployment at scale.
We are looking for a highly motivated Summer Intern to work directly on the design, implementation, and optimization of Reinforcement Learning (RL) environments for Physical AI and Vision-Language Models (VLMs / VLA). You will help create high-fidelity, multi-modal simulation environments that power our post-training and data flywheel. This is a hands-on role with real ownership — your environments will be used by our Nvex orchestration layer and customer models.
What You'll Do
- Design and implement domain-specific RL environments for Physical AI tasks (robot manipulation, locomotion, dexterous grasping, human-robot interaction, etc.).
- Build and extend multi-modal observation spaces (RGB + depth + tactile + proprioception + language instructions) using modern simulators such as Isaac Lab, Isaac Gym, MuJoCo, SAPIEN, or custom environments.
- Develop reward functions, termination conditions, and curriculum learning pipelines that align with real-world robot behavior and VLM/VLA objectives.
- Integrate RL environments with our post-training stack (SFT → RLHF/PPO/DPO-style training loops) and evaluation harness.
- Run large-scale experiments, benchmark environment fidelity (sim-to-real gap), and iterate rapidly based on model performance feedback.
- Collaborate closely with ML engineers and the founding team on production-grade code and documentation.
Requirements
- Currently pursuing a Master's or PhD degree in Computer Science, Robotics, AI/ML, or a related field (graduating 2026 or later).
- Strong proficiency in Python and PyTorch.
- Hands-on experience with at least one RL framework (Gymnasium, Stable-Baselines3, RLlib, or CleanRL).
- Familiarity with at least one robotics simulator (Isaac Lab/Gym, MuJoCo, PyBullet, etc.).
- Solid understanding of RL fundamentals (MDPs, policy gradients, value functions, PPO/SAC, etc.).
Nice to Have
What You'll Gain
- Ownership of production RL environments that ship to real customers.
- Exposure to the full stack: data pipeline → RL training → evaluation → agent orchestration.
- A fast-paced startup environment where your work directly impacts the next wave of Physical AI.
- Competitive summer stipend + potential for full-time offer or strong recommendation letters.
About the Role
Alchedata is building the self-improving data layer for Physical AI. We focus on a problem beyond "more data" or "bigger models": when a VLM, world model, VLA system, or embodied agent fails in the real world, how do we systematically convert that failure into better evaluation, better data decisions, better simulation, and ultimately a better checkpoint?
Our core loop is: Failure → Evaluation → Data Recipe → Data Operation / Simulation → Better Checkpoint. We are not building a generic data platform. We are building the infrastructure that helps Physical AI systems improve continuously after deployment.
What This Role Is About
Most teams already have some combination of training infrastructure, data storage, annotation workflows, simulation tooling, and evaluation scripts. What is usually missing is the layer that connects them.
- What exactly failed?
- What evidence is missing?
- What should be collected, curated, generated, or replayed next?
- How should synthetic and real-world data be combined?
- How do we validate that the next checkpoint is actually better without breaking regression gates?
This role exists to design the data architecture, object model, evidence chain, metadata, lineage, and governance layer behind those decisions.
What You'll Do
- Design a unified data architecture for VLM / world model / VLA workflows spanning multimodal raw data, evaluation outputs, failure taxonomies, DataRecipe objects, simulation environments, replay evidence, and checkpoint lineage.
- Build the evidence chain behind failure diagnosis → data decision → simulation validation → checkpoint promotion.
- Define durable schemas, metadata systems, ontologies, lineage models, and access patterns so data becomes diagnosable, traceable, reusable, and verifiable.
- Work closely with research, ML, simulation, data ops, and platform engineering teams to ensure the architecture directly supports model improvement rather than generic storage.
- Build core infrastructure for failure clustering, root-cause analysis, targeted data collection and curation, real + synthetic data mix management, eval suite generation, environment packaging, replay validation, and checkpoint regression tracking.
- Help define foundational platform objects such as EvalReport, DataRecipe, EnvironmentPackage, Replay Evidence, and Checkpoint Comparison / Promotion Records.
- Establish strong governance boundaries across customer separation, access control, anonymization, and aggregated insight reuse.
- Make sound architectural tradeoffs in a fast-moving startup while still building for long-term leverage.
Requirements
- BS/MS or equivalent practical experience in Computer Science, Data Engineering, ML, Robotics, Systems, or a related field.
- Strong experience in multimodal data platforms, ML data infrastructure, dataset lifecycle systems, robotics or embodied AI data pipelines, simulation data pipelines, experiment tracking, model lineage, or evaluation infrastructure.
- Strong understanding of the ML data lifecycle, including collection, annotation, QA, slicing, versioning, replay, archival, and reuse.
- Ability to independently design complex schemas, metadata systems, ontologies, lineage models, and access models.
- Familiarity with the data requirements of VLMs, world models, VLA systems, or embodied AI, including temporal consistency, multimodal alignment, success/failure/subtask labeling, sim-to-real coordination, and governance of mixed real and synthetic data.
- Strong hands-on ability with Python, SQL, and at least one large-scale data processing or storage stack.
- Strong cross-functional communication and execution skills.
Nice to Have
Why This Role Matters
This is not a support function. This is a foundational role in how Physical AI systems improve over time. The architecture you design will directly influence how quickly failures are diagnosed, how precisely the next data decision is made, how tightly evaluation, simulation, and checkpoint evolution are connected, and how today's workflows become tomorrow's product moat.
What You'll Gain
- Work on foundational infrastructure for self-improving Physical AI, not generic data plumbing.
- Own design decisions that directly affect model improvement loops.
- Partner closely with research, platform, simulation, and product leaders.
- Build from 0 to 1 in a high-agency environment with meaningful technical depth and product impact.
About the Role
Alchedata is building the self-improving data layer for Physical AI. We focus on a problem beyond simply training bigger models: when a VLM, world model, VLA system, or embodied agent fails in the real world, how do we systematically turn that failure into the next verifiable model improvement cycle?
Our core loop is: Failure → Evaluation → Data Recipe → Data Operation / Simulation → Better Checkpoint. In that loop, reward models, evaluators, and verifiers are not just scoring tools. They are foundational infrastructure for continuous improvement.
What This Role Is About
Many teams can train models and run benchmarks. The harder challenge starts when models enter real environments and encounter long-tail failures, partial success, unstable gains, and ambiguous signals. The real work is being able to answer:
- Did the model truly succeed or fail, and at what step or subtask?
- Which reward, judge, or verifier signals actually help models improve rather than merely score well on benchmarks?
- How do we let systems learn continuously from rollouts, replay, simulation, and historical experience?
- How do we validate that a new checkpoint is genuinely better without reward hacking or narrow-metric overfitting?
This role exists to turn those questions into trainable, deployable, and scalable reward modeling and evaluation systems.
What You'll Do
- Design and train reward models, judge models, verifiers, and success predictors for VLM, world model, and VLA settings.
- Build task-level, subtask-level, and temporal reward systems for success/failure detection, progress estimation, subtask completion, confidence calibration, and temporal failure localization.
- Develop world-model-based approaches for imagined rollout evaluation, counterfactual analysis, trajectory scoring, and policy improvement.
- Build learning loops from real-world failures: rollout → evaluation → diagnosis → recipe / improvement signal → checkpoint validation.
- Work closely with data, simulation, platform, and robotics / embodied AI teams to integrate reward and evaluation modules into training, replay, regression testing, and checkpoint promotion workflows.
- Advance areas such as preference learning, pairwise ranking, process reward models, outcome reward models, multimodal verifiers, model-based RL, test-time adaptation, failure clustering, and root-cause attribution.
- Design robust benchmark suites and offline / online evaluation protocols to improve reliability, generalization, and interpretability.
- Produce strong technical outputs, internal benchmarks, production implementations, and publications or patents where appropriate.
Requirements
- MS/PhD in Computer Science, Machine Learning, Robotics, Statistics, Applied Math, or a related field, or equivalent practical experience.
- Strong experience in reward modeling, RL / RLHF / preference learning, multimodal learning / VLM, world models / model-based learning, VLA / embodied AI, or robot learning.
- Strong hands-on experience with deep learning frameworks such as PyTorch or JAX, and the ability to build training and evaluation pipelines independently.
- Strong understanding of reward hacking, distribution shift, calibration, OOD generalization, and evaluation leakage.
- Strong problem formulation and experimental design skills; able to turn messy real-world failures into rigorous hypotheses and validation frameworks.
- Strong publication, implementation, and cross-functional collaboration skills.
Nice to Have
Why This Role Matters
This is not a narrow offline scoring role. It is a foundational role in the continuous improvement loop for Physical AI. Every reward signal, judge logic, evaluation protocol, and validation pipeline you build can directly influence whether systems truly learn from real-world failure.
What You'll Gain
- Work on foundational infrastructure for self-improving Physical AI rather than isolated benchmark optimization.
- Own systems that directly affect model improvement loops and checkpoint quality.
- Partner closely with research, platform, simulation, and product leaders.
- Build from 0 to 1 in a high-agency environment with meaningful technical depth and product impact.
About the Role
Alchedata is building the self-improving data layer for Physical AI. Most Physical AI teams already have pieces of the stack: training infrastructure, data storage, evaluation scripts, simulation tools, and deployment monitoring. What is still missing is the layer that connects real-world failure to the next improvement cycle.
That is what Nvex is designed to do. Within Nvex, AlcheSim is the environments module. It builds and validates replayable simulation environments from text, image, video, and verified assets; supports cross-simulator transfer; and produces the replay evidence used in checkpoint validation and promotion.
What This Role Is About
This is a simulation engine-building role, not a simulation-usage role.
You will be a core engineer on AlcheSim, responsible for turning an already working environment-generation and replay pipeline into a scalable product capability. Your job is to help make simulation a living part of the model-improvement loop: when a model fails in the real world, the system should be able to generate or refine the right environment, replay the failure, validate candidate fixes, and feed that evidence into checkpoint comparison.
In practical terms, you will help build the infrastructure that connects real-world failure → environment generation → replay → validation → better checkpoint.
What You'll Do
- Build the core AlcheSim engine for environment and scene generation from multimodal inputs such as text, image, video, and verified assets.
- Develop environment assembly pipelines that are physically coherent, reusable, and suitable for replay and validation workflows.
- Connect failure diagnosis to environment generation, so evaluation results and real-world failure cases can trigger targeted environments where models are currently weakest.
- Build replay and validation infrastructure including deterministic replay, sim-consistency checks, regression gates, and comparison tooling, so simulation outputs can be trusted as evidence for checkpoint promotion.
- Enable cross-simulator portability so environments and scenarios can move across engines such as Isaac Sim / Isaac Lab, MuJoCo, and other backends rather than being locked to one stack.
- Productize internal workflows into robust, versioned, self-serve-ready platform capabilities: environment packaging, cataloging, APIs, CI-integrated scenario libraries, and cloud-scale execution.
- Work closely with the research team to translate self-evolving learning mechanisms, such as world-model rollouts, test-time adaptation, and online policy improvement, into simulation engine capabilities.
- Partner with evaluation and data operations teams so simulation, real-world collection, and evaluation form one closed loop rather than separate workflows.
Requirements
- 3+ years building simulation systems or simulation infrastructure in robotics, embodied AI, autonomous driving, game engines, or physics-based environments.
- Experience building engines, pipelines, or platform infrastructure, not only using simulation tools.
- Hands-on experience with closed-loop simulation systems, replay / log-driven simulation workflows, large-scale scenario or environment generation, or sim-to-real / sim-consistency evaluation.
- Strong software engineering fundamentals, including Python and C++ or Rust, Linux, Docker, Git, and CI/CD.
- Experience running simulation workloads at scale on cloud infrastructure, clusters, or distributed compute environments.
- Experience building data or evaluation closed loops, such as turning field failures, evaluation outputs, or regression findings into new scenarios, test suites, or environments in an automated or semi-automated way.
- Ability to work across the stack, from simulation and systems details through product APIs, packaging, and user workflows.
Nice to Have
Why This Role Matters
You will own a core part of AlcheSim, the environments module of Nvex, and help define how replayable environments become part of a broader Physical AI improvement loop. This is not about building prettier scenes. It is about making simulation operationally useful: every real-world failure should be able to produce better replay environments, better validation, and ultimately a better checkpoint.
What You'll Gain
- Own a core part of AlcheSim and help define how simulation becomes part of the model-improvement loop.
- Work closely with a research team focused on self-evolving robot learning and production-capable systems.
- Shape engine architecture, replay infrastructure, cross-simulator portability, and productization.
- Build a category-defining platform with flexible location across Bay Area, Singapore, or remote.