D1 Available for Acquisition and Licensing

Private coding-agent tasks without the production wait

Evaluate a substantial collection of unpublished terminal and repository-level software-engineering tasks for model post-training, reinforcement learning and private evaluation programs.

Private · Unpublished · Sample Access Available
Private Coding-Agent Tasks
4,500+
Terminal-Bench Tasks
2,000
Terminal-Bench Science Tasks
1,000
SWE-bench Tasks
1,500
DATASET OVERVIEW

Three collections, ready to inspect.

Private task inventory built to the exact formats used by the industry's best-known coding-agent benchmarks.

Terminal-Bench
2,000 Tasks Available

Deaimer's private task collection built to the Terminal-Bench task and environment format — difficult, multi-step terminal workflows with objectively verifiable outcomes, covering system administration, DevOps, debugging, security and data processing.

Task Count2,000 private tasks
DomainTerminal operations, Linux & sysadmin, DevOps, debugging, security, data processing
FormatTerminal-Bench-compatible
AvailabilitySample access by request
Terminal-Bench Science
1,000 Tasks Available

A private task collection built to the Terminal-Bench Science format — terminal-driven scientific-computing and research workflows, from data analysis pipelines to computational-methods tasks with objectively verifiable outcomes.

Task Count1,000 private tasks
DomainScientific computing, data analysis, research & simulation workflows
FormatTerminal-Bench Science-compatible
AvailabilitySample access by request
SWE-bench
1,500 Tasks Available

A private task collection built to the SWE-bench repository format — repository-level software-engineering problems requiring codebase understanding, multi-file changes, and test-driven resolution.

Task Count1,500 private tasks
DomainRepository bug fixing, feature implementation, test repair, multi-file changes
FormatSWE-bench-style
AvailabilitySample access by request

Deaimer is an independent contributor building task collections compatible with these benchmark formats. Deaimer does not own, operate, sponsor or hold an official partnership with Terminal-Bench or SWE-bench, and does not display their marks without permission.

Need a task set that doesn't exist yet? See Custom RL Environments. Need to validate model performance against a benchmark? See Benchmarks & Research.

PACKAGE CONTENTS

What a delivery can include.

Available package contents vary by collection and can include:

Task instructions

Task prompts and setup instructions provided to the agent.

Task metadata & taxonomy

Structured labels for category, difficulty, and capability area.

Containerized environments

Environment definitions used to run each task reproducibly.

Source repositories

Underlying repository state for repository-level tasks.

Fixtures & artifacts

Supporting data and files required to run a task.

Automated graders

Scoring logic used to verify task completion.

Test suites

Test coverage associated with repository-level tasks.

Reference solutions

A working solution used to validate solvability.

Oracle results

Recorded results from a known-correct solution run.

Negative-control results

Recorded results from a known-failing baseline run.

QA records

Documentation from the independent review process.

Model-run evidence & agent trajectories

Trial run data and agent trajectories captured during validation.

Difficulty & capability classifications

Classification of each task by difficulty and capability area.

Not every item is included in every collection — availability is confirmed during the sample review.

USE CASES

Where this data gets used.

Coding-agent post-training

Fine-tune coding agents against high-difficulty terminal and repository tasks.

Supervised fine-tuning

Use verified task and solution pairs as supervised training data.

RLVR & reinforcement learning

Verifiable-outcome tasks suited to RLVR and RL training pipelines.

Private model evaluation

Run confidential evaluations that avoid public benchmark contamination.

Capability-gap analysis

Identify where a model's coding-agent capabilities fall short.

Grader & reward-model research

Develop and validate graders and reward models against varied tasks.

Agent robustness testing

Probe agent behavior against difficult, adversarial-adjacent tasks.

Benchmark extension

Extend an existing benchmark's coverage with additional private tasks.

Failure-mode analysis

Study how and where agents fail across a broad task distribution.

LICENSING OPTIONS

Ways to acquire access.

Non-Exclusive Licence

Use the dataset while Deaimer retains the ability to license it to additional customers.

Exclusive Licence

Negotiate exclusive rights for a defined collection, capability area, language, format, market, or period.

Direct Acquisition

Acquire an agreed dataset package and associated rights under a negotiated asset-purchase or data-licensing agreement.

Final scope, permitted uses, exclusivity, warranties, delivery formats and intellectual-property rights are defined contractually.

EVALUATION PROCESS

From inquiry to delivered package.

1
Submit an inquiry

Tell us about your intended use and the collection you're interested in.

2
Execute an NDA

A mutual NDA is put in place before any detailed material is shared.

3
Review the catalog

Review a catalog and metadata summary describing available tasks.

4
Inspect sample tasks

Review a small set of controlled sample tasks from the collection.

5
Diligence & overlap checks

Agree on technical diligence and dataset-overlap checks with your team.

6
Select a structure

Choose a licensing or acquisition structure that fits your use case.

7
Receive the package

Receive the versioned delivery package under the agreed terms.

SAMPLE ACCESS

Inspect the data before discussing a larger transaction.

Request a structured catalog and a small set of controlled samples under your preferred confidentiality process.