About · DTRM Research Lab
Research by building. Judgment by design.
Human Clause is the public flagship of DTRM Research Lab: an independent practical AI research initiative built to turn hypotheses into code, experiments and measurable evidence while keeping scientific acceptance and responsibility human.
Urtzi Arana
Urtzi Arana Santamaria is a digital product leader and independent researcher working at the intersection of payments, computational mathematics and artificial intelligence. His professional background includes building and scaling digital payment products in banking and fintech environments, with a strong focus on customer value, risk, engineering and execution.
His academic and research path connects business, Computational Mathematics, AI decision systems, game theory, quantitative finance and continued study of quantum information and technologies. That combination shapes a product-minded research approach: build the system, stress it, measure it and then decide what the evidence actually supports.
He founded DTRM Research Lab to make that process explicit and reproducible — and to test a recurring question: as machines become more capable, where does human judgment still add measurable value?
Public LinkedIn profile →Why DTRM Research Lab exists
The lab began by turning a working investment experiment into a reproducible laboratory for AI, decision science and optimization.
Research by building. Every hypothesis should become code, an experiment and a measurable result.
Decision quality over hype. Machine-only, human-guided and hybrid systems are compared under the same rules.
Safety through guardrails. Reproducibility, leakage, costs and out-of-sample evidence matter before promotion.
Publish what survives. The objective is a growing body of practical evidence, not a collection of demos.
The operating system
Create. Test. Doubt. Explain.
DTRM Research Lab follows a Feynman-inspired research method: understand by creating, distrust certainty and let experiment decide. A model that cannot be reproduced, challenged or explained is not treated as a result simply because its headline metric improved.
01
Create
Build the model, pipeline and optimizer ourselves.
02
Test
Hold out data, compare with frozen baselines and reproduce results.
03
Doubt
Search actively for leakage, regime dependence and false improvements.
04
Explain
Make the evidence understandable enough to challenge and repeat.
Agentic Graph Engineering
Agentic Graph Engineering is the lab's way of using AI agents to accelerate implementation without allowing probabilistic systems to become the final judge of scientific truth.
Human hypothesis
Define the research question and non-negotiable rules.
Agents propose
Architect, code, test, review and document in parallel.
Gates verify
Deterministic tests, leakage checks, backtests and thresholds decide pass or fail.
Human approves
Merge or promote only after evidence and reviewer output are visible.
LLMs may help decide what to try next. They do not decide whether the science passed.
The engineering graph
Specialist roles are separated so implementation is not its own validator. The graph can route work across research and engineering roles while converging through the same CI and promotion gates.
Developer ≠ Validator ≠ Risk reviewer.
DTRM AI Agent · OpenAI
The DTRM AI Agent is the lab's research and engineering collaborator built with OpenAI capabilities. It helps search the problem space, inspect repositories, propose architectures, implement code, design tests, review evidence and document decisions at machine speed.
Its authority deliberately stops before scientific acceptance. The agent cannot rewrite protected stage rules, declare its own experiment successful or bypass deterministic gates. Human approval remains the final promotion boundary.
OpenAI provides underlying AI capabilities used by the research agent. DTRM Research Lab and Human Clause are independent projects and are not publications or endorsements by OpenAI.
Stage contracts: rules agents cannot negotiate
Each research stage receives an explicit specification: parent baseline, invariants, required tests, output contracts and promotion gates. The request is never simply “improve the model”.
Data splits, embargo rules, benchmark definitions, transaction-cost policies and risk thresholds remain protected. An agent can argue for a change in a future experiment; it cannot silently change the current game to make itself win.
The model can argue. The gate cannot.
What the lab is really building
The first portfolio is a laboratory. The durable asset is the research system and the evidence it produces.
Reproducible research engine
Versioned data, configurations, scenarios and experiment history.
Scenario library
A growing catalogue of adversarial regimes and model-failure patterns.
Human-judgment dataset
A record of when humans override machines — and whether the override helped.
Cross-solver benchmark
Classical, quantum-inspired, hybrid and gate-based optimization on comparable problems.
Publishable evidence
Cohort-by-cohort results that can support papers, technical notes and open research.
Reusable agentic harness
A governed way to build new research systems faster without losing auditability.
The competitive advantage is not one model. It is the ability to generate trustworthy model improvements repeatedly.
