Executable environments
Resettable, observable worlds for training and evaluating agents across repositories, terminals, and browser-based workflows.
Independent AI research · San Francisco + Chicago
We study reinforcement-learning environments, planning, and reliable agent systems for software, web, terminal, and operational work.
Better agents require more than better policies. They need worlds with legible state, meaningful consequences, trustworthy feedback, and enough variation to learn skills that transfer.
Resettable, observable worlds for training and evaluating agents across repositories, terminals, and browser-based workflows.
Testable outcome contracts, partial-credit signals, and adversarial checks that distinguish task completion from reward hacking.
Task generation and curricula that transfer skills from shell operations to software engineering and stateful web work.
Reasoning, search, and recovery for long-horizon agents operating with limited context, time, permissions, and evidence.
Memory and state architectures that make long-running work resumable, auditable, and robust to interruption.
Applying reliable agent techniques to document operations, robotics, manufacturing, procurement, and finance.
A practical research framework for building RL environments across software engineering, terminal, and web tasks—where every episode can be reset, every outcome can be verified, and every failure produces useful evidence.
STATE_01InitializeMaterialize a controlled, inspectable task state.ACT_02InteractExpose consequential actions through real tools.VERIFY_03EvaluateCheck outcomes against independent reward contracts.RESET_04RecoverCapture evidence, reset cheaply, and vary the next episode.Products, prototypes, and research systems from Neural Intelligence Labs. Each one starts with the same question: what would make an intelligent system more capable, legible, and trustworthy?

Agents that can pick up where they left off.
A stateful runtime for long-running agent work—keeping plans, tool actions, evidence, checkpoints, and recovery separate from the conversation window.
The idea: an agent should survive interruption without losing the plot.
Case study in progress
Find the fault. Make the smallest repair. Prove it.
A repository-repair workflow that turns debugging into an evidence loop: reproduce one failure, isolate the layer, propose a focused patch, and run the decisive test.
The idea: a useful coding agent returns a verified repair, not a plausible diff.
Case study in progress
Documents in. Reconciled action out.
An AI document-operations workspace for reviewing invoices and inbox documents, reconciling records, routing approvals, and syncing clean evidence into business systems.
The idea: automation becomes trustworthy when every action carries its proof.
Explore product
Give software agents a world they can learn from.
A research build for resettable, verifiable training environments spanning terminal, repository, and browser tasks—with state, consequences, rewards, and recovery.
The idea: environment quality sets the ceiling on what an agent can learn.
Case study in progress
Protect attention long enough to do work that matters.
A calm focus app in development for turning intention into a protected work session—reducing the pull of fragmented attention and making sustained progress visible.
The idea: focus should feel like a space you enter, not a punishment you endure.
Case study in progress
Turn a look you love into something you can actually find.
A prompt-to-style shopping app that uses text or visual inspiration to discover fashion across multiple stores, personalize recommendations, and preview choices with virtual try-on.
The idea: shopping should begin with your taste—not a retailer’s endless feed.
Explore productIndex turns invoices, inbox documents, approvals, and posting evidence into one review-to-sync workspace. It is the document layer between messy operational work and the systems clients already run.
Discuss IndexTie every mismatch to its source document, proposed owner, and approval path.
Turn loose documents into tagged, assigned, reviewable work.
Show GL, cash, inventory, risk, and source evidence before anything syncs.
Connect orders, invoices, payments, and notes to the missing next step.
Link counts, shortages, and reorder triggers to the documents behind them.
Keep open AR, approvals, and blocked work visible in one cash queue.
Map document sets, client systems, and exception rates.
Define blueprint fields, validators, queues, and evidence-backed review states.
Run partner-led acceptance with evidence attached to every extracted value.
Sync approved work into the ERP, TMS, CRM, or vertical system of record.
Connects with Salesforce, QuickBooks, HubSpot, Stripe, NetSuite, Anrok, DocuSign, Gmail, Mercury, Plaid, Slack, and the vertical systems clients already depend on.
For ERP and accounting consultants configuring Index inside client rollouts.
For firms building a repeatable document-operations practice with Index.
For software platforms embedding Index workflows inside their own product.
Start with one workflow
Begin with invoice match, customer follow up, posting proof, or the cash dashboard—then expand after the operating path is working.
One visual language, many systems. The registration mark and chartreuse signal identify work made inside Neural Intelligence Labs—before you read the name.

Founder & researcher
Reinforcement learning · Planning · Agent systems
Mehrdad is a computer scientist and founder of Neural Intelligence Labs. His work connects reinforcement learning, planning, and human-centered agent systems with the systems problems that make long-running AI work reliable in practice.
Open to research collaborations, benchmark partnerships, visiting talks, and working sessions in San Francisco + Chicago around agent environments and reliable AI systems.
Start a conversation