Fractional contractNew York City preferred · remote possible
Fractional AI Security Researcher
Turn agent-security hypotheses into reproducible evidence, evaluations, and product direction.
Why this role
Thoth needs rigorous research on model behavior, agent attacks, policy verification, and security evaluation.
This role pairs research depth with practical prototypes and explicit production handoffs rather than open-ended exploration.
How we hire
High Energy. High Intelligence. Low Ego.
H₂O is Aten's standard for enthusiasm, competence, and character. We look for people who create momentum, exercise sound judgment, and put truth and the mission ahead of credit.
Design and evaluate probabilistic verifiers that can only make deterministic policy decisions more cautious.
Build reproducible attacks and evaluations for prompt injection, tool poisoning, identity spoofing, model misuse, adversarial manipulation, and multi-agent failures.
Create versioned datasets, experiment manifests, ablations, scorecards, and regression thresholds.
Evaluate quality, latency, reliability, scalability, and cost rather than optimizing a benchmark in isolation.
Prototype instruction-provenance and independent-action-attestation techniques.
Contribute experiments, limitations, and reproducibility material to Aten research publications.
Expected outcomes
Milestones, not activity theater.
First week
Reproduce a baseline and submit a reviewable experiment, fixture, test, or research correction within three to five business days.
Initial engagement
Deliver a written hypothesis and threat model, reproducible code, versioned inputs, defined metrics, results, limitations, and a production handoff recommendation.
What we are looking for
Demonstrated depth through a PhD, published research, offensive security work, or production AI-security systems.
Strong Python and modern machine-learning framework experience.
Experience implementing evaluation harnesses with versioned datasets, reproducible execution, grader validation, and failure analysis.
Strong offensive or application security foundations and the ability to build safe proofs of concept.
Clear technical writing and the discipline to report negative or inconclusive findings.
Ability to translate research into defensive controls and production requirements without overstating results.
Useful experience
Go or Rust prototyping
Agent frameworks, MCP, adversarial ML, or model evaluation
Bug bounty, red-team, military, intelligence-community, or high-value-target research
Technical publications, conference talks, or open-source security tools
Our working stack
Python model, verifier, and evaluation services
Go product APIs and control plane
Rust endpoint and runtime services
AWS and Terraform-managed research environments
How we evaluate
We use a hybrid process because the job requires both independent engineering judgment and effective use of AI. You should be able to reason without a model, then use one to move faster while reviewing its output critically.
1A focused introduction call about the engagement and your relevant work.
2A practical discussion or scoped exercise based on a representative deliverable.
3Agreement on outcomes, access boundaries, availability, and initial engagement timing.
Location and travel
New York City preferred · remote possible
Occasional team or research-program travel may be requested.