EngineeringThe AI Agent Instrumentation Tax: Lessons from 1,000 Hours of Runtime Telemetry StagingRead the post →
← All roles
Fractional contractNew York City preferred · remote possible

Fractional AI Security Researcher

Turn agent-security hypotheses into reproducible evidence, evaluations, and product direction.

Thoth needs rigorous research on model behavior, agent attacks, policy verification, and security evaluation.

This role pairs research depth with practical prototypes and explicit production handoffs rather than open-ended exploration.

How we hire

High Energy. High Intelligence. Low Ego.

H₂O is Aten's standard for enthusiasm, competence, and character. We look for people who create momentum, exercise sound judgment, and put truth and the mission ahead of credit.

Read our hiring philosophy

What you will own

  • Design and evaluate probabilistic verifiers that can only make deterministic policy decisions more cautious.
  • Build reproducible attacks and evaluations for prompt injection, tool poisoning, identity spoofing, model misuse, adversarial manipulation, and multi-agent failures.
  • Create versioned datasets, experiment manifests, ablations, scorecards, and regression thresholds.
  • Evaluate quality, latency, reliability, scalability, and cost rather than optimizing a benchmark in isolation.
  • Prototype instruction-provenance and independent-action-attestation techniques.
  • Contribute experiments, limitations, and reproducibility material to Aten research publications.

Milestones, not activity theater.

First week

Reproduce a baseline and submit a reviewable experiment, fixture, test, or research correction within three to five business days.

Initial engagement

Deliver a written hypothesis and threat model, reproducible code, versioned inputs, defined metrics, results, limitations, and a production handoff recommendation.

What we are looking for

  • Demonstrated depth through a PhD, published research, offensive security work, or production AI-security systems.
  • Strong Python and modern machine-learning framework experience.
  • Experience implementing evaluation harnesses with versioned datasets, reproducible execution, grader validation, and failure analysis.
  • Strong offensive or application security foundations and the ability to build safe proofs of concept.
  • Clear technical writing and the discipline to report negative or inconclusive findings.
  • Ability to translate research into defensive controls and production requirements without overstating results.

Useful experience

  • Go or Rust prototyping
  • Agent frameworks, MCP, adversarial ML, or model evaluation
  • Bug bounty, red-team, military, intelligence-community, or high-value-target research
  • Technical publications, conference talks, or open-source security tools

Our working stack

  • Python model, verifier, and evaluation services
  • Go product APIs and control plane
  • Rust endpoint and runtime services
  • AWS and Terraform-managed research environments

How we evaluate

We use a hybrid process because the job requires both independent engineering judgment and effective use of AI. You should be able to reason without a model, then use one to move faster while reviewing its output critically.

  1. 1A focused introduction call about the engagement and your relevant work.
  2. 2A practical discussion or scoped exercise based on a representative deliverable.
  3. 3Agreement on outcomes, access boundaries, availability, and initial engagement timing.

Location and travel

New York City preferred · remote possible

Occasional team or research-program travel may be requested.