Aniket Mallick — headshot
Aniket Mallick — Bengaluru, India

I build AI you can refuse.

Robot arms with trust gates in the control code — and a verification layer for the data robots learn from. Everything below is a ledger entry: dated, linked, or labeled private. Nothing embellished.

Toceta is the practice that came out of it: pre-registered evals for robot learning policies — the pass/fail bar written in public before the robot moves, the result published by the date, either way. The first one, run on my own arm as the method's own test: PR-001 → CERT-001. Independence is the next thing to earn.

Watch: the trust layer, off and on, in 53 seconds.

2 products acquired ex-C3 AI ex-Oracle ACM-ICPC regionals

Evaluating me for a program? The 53-second demo is the fastest read — then my inbox is open.

Now — the experiment

Current obsession: can a robot's claim be checked by someone who didn't build it? I wrote down, in public, what one robot task would be judged on — before the robot moved — and published the result by the date, pass or fail. The first one is done, on my own arm, by me; the document says so. The next pre-registration (PR-002, October) tests a learned policy against a harder bar. The next piece is the layer between a learned policy and a human — built so every check is seen to fire and can be read back afterwards.

2026-10
  1. 2026-10-04
    The same arm run with a runtime trust layer off, then on: a hand monitor and heartbeat interlock that freeze the arm within one frame, a vision-model judge asked at checkpoints whether a handover is safe (answer and latency logged), re-orientation of the pointed end before approach, placement in the open palm, and a hash-chained step log replayed as the right-hand panel of the video. A demonstration on an SO-101 with a real screwdriver and my own hand in both runs; not part of any pre-registered evaluation.
2026-09
  1. 2026-09-29
    Certificate of Conformance to Pre-Registration PR-001 published, one day early. Policy of record: the pre-registered scripted baseline (no camera, no learning — the readiness gate for the learned policy failed by the clock, erratum 5). Result of record: 29 of 40, Wilson 95 % [57 %, 84 %], against a bar of ≥ 22/40 written for a learned policy. The finding: the task as registered is too easy to tell a learned policy from a script; PR-002's bar goes above 84 %. Nine dated errata and a base guard that couldn't measure its own bar are in §2 of the certificate, before the result; one trial corrected against its own photo is in §3, beside it.
  2. 2026-09-25
    PR-001 posted: the public pre-registration — task, bars, tails, exclusion rules, what is not covered — frozen at tag prereg-001. Planned for 10 Sep; the slip is stated in the document, not hidden.
    PR-001 SHA-256 81f371b2…
  3. 2026-09-24
    The robot base was found unfixed to the table: it had turned about 4° in 22 hours between two registrations. Clamped, re-registered, disclosed. Whether it explains a change to this rig observed before the anchor session cannot be said — the one record that could have dated that onset was overwritten (PR-001 §9).
  4. 2026-09-18
    A headline finding of my own (“coverage predicts capability”) did not survive measurement: nearest-demo distance 4.9° for successes and 4.9° for failures. Downgraded to observed, untested and published as a dated correction.
2026-08
  1. 2026-08-16
    First result: 50 self-recorded demos → 28% full completion, 44% grasp, 78% contact on a $200 arm — vs 0 baseline. Three of five completions via emergent recovery.
  2. Fine-tuned a 450M-parameter VLA (SmolVLA) on my own 50-episode dataset — overnight, through three Colab crashes.
  3. Hardened the eval harness: selectable clamp profiles, checkpoint-native normalization, self-describing results file — every trial row carries model ref, revision, and achieved control rate. Provenance is the product.
  4. Recorded 50 clean teleoperated demonstrations of one pick task across a marked 5×5 grid — the teleop control arm of the experiment, collected by hand in one night.
2026-07
  1. Zero-shot baseline measured: an untuned generalist VLA on this arm scored 0/2. The floor is real, and it's mine to beat.
  2. Built ARM-ANI in an 18-hour hackathon: voice-interactive robot arm with trust gates in the control path — asks when ambiguous, states confidence aloud, waits for approval below threshold, stands down after 10 seconds of silence. Every decision goes to an audit log. Prompts are style; gates are law.

Next: PR-002 — a learned policy on the same arm, same task, bar set above the null's upper confidence bound, every gate designed before the first motion. The egocentric-video question is capped for now; its instrument became the audit practice.

Shipped and acquired

Products that left my hands.

2026 Medical-AI CRM

Demand forecasting for medicines. Cumulative demand views and outlier detection on medicine demand trends, built for the medical sector. Acquired. Buyer and terms private.

acquired · details private

2026 Patient platform on WhatsApp

End-to-end: patients send documents and questions over WhatsApp; the system ingests records and returns AI-generated diet and care responses. Acquired. Buyer and terms private.

acquired · details private

Await Arcade

Atomic cognitive games for the seconds your agent is thinking — play while Claude Code runs, switch back when it's done. Used in hackathons for collaborative sessions and competitions.

TRIBE v2 Creative Workbench

Compare how two video stimuli activate the brain — an open-source workbench on Meta's TRIBE v2 foundation model. Predicts cortical activity across 20,484 brain vertices per video and renders interactive, frame-by-frame difference maps. GPU-free demo mode; MIT-licensed.

Prompt-to-Product — hardware

Prompt to Physical Product.

in progress · not yet public

Life-OS

This site's previous life. An AI that graded my discipline in public: tamper-evident, SHA-256 hash-chained, publicly verifiable. Frozen at entry #113 · 2026-10-04 — the robot ate the roadmap, and the freeze is itself verifiable.

Track record

C3 AI (2024–2026) — built the glass-box explainability module for enterprise demand forecasting: planners see why the model decided, and can override it. 33% stockout reduction · £542K revenue impact · forecasting workflows over 1M+ SKU-location subjects. Highest performance rating within 6 months.

Oracle (2022–2024) — Java REST APIs for the NFVD orchestrator behind telecom systems serving 1B+ users: +40% lifecycle efficiency, −37% production tickets, CI/CD security automation that cut manual effort 80%. 2× quarterly Development Excellence Award, SVP-nominated.

ACM-ICPC regionals (2019–20) · B.Tech CSE, IEM Kolkata · CGPA 9.14.

Community

Off the keyboard, I advise builder communities on designing and running hackathons — formats, judging, operations. Await Arcade shows up at those events too, as a collaborative-session and competition layer. Before any of that: designed and taught a data-structures & algorithms curriculum to US graduate and career-transition students — 70% course completion-to-conversion.