Private coding-agent tasks without the production wait
Evaluate a substantial collection of unpublished terminal and repository-level software-engineering tasks for model post-training, reinforcement learning and private evaluation programs.
Three collections, ready to inspect.
Private task inventory built to the exact formats used by the industry's best-known coding-agent benchmarks.
Deaimer's private task collection built to the Terminal-Bench task and environment format — difficult, multi-step terminal workflows with objectively verifiable outcomes, covering system administration, DevOps, debugging, security and data processing.
A private task collection built to the Terminal-Bench Science format — terminal-driven scientific-computing and research workflows, from data analysis pipelines to computational-methods tasks with objectively verifiable outcomes.
A private task collection built to the SWE-bench repository format — repository-level software-engineering problems requiring codebase understanding, multi-file changes, and test-driven resolution.
Deaimer is an independent contributor building task collections compatible with these benchmark formats. Deaimer does not own, operate, sponsor or hold an official partnership with Terminal-Bench or SWE-bench, and does not display their marks without permission.
Need a task set that doesn't exist yet? See Custom RL Environments. Need to validate model performance against a benchmark? See Benchmarks & Research.
What a delivery can include.
Available package contents vary by collection and can include:
Task instructions
Task prompts and setup instructions provided to the agent.
Task metadata & taxonomy
Structured labels for category, difficulty, and capability area.
Containerized environments
Environment definitions used to run each task reproducibly.
Source repositories
Underlying repository state for repository-level tasks.
Fixtures & artifacts
Supporting data and files required to run a task.
Automated graders
Scoring logic used to verify task completion.
Test suites
Test coverage associated with repository-level tasks.
Reference solutions
A working solution used to validate solvability.
Oracle results
Recorded results from a known-correct solution run.
Negative-control results
Recorded results from a known-failing baseline run.
QA records
Documentation from the independent review process.
Model-run evidence & agent trajectories
Trial run data and agent trajectories captured during validation.
Difficulty & capability classifications
Classification of each task by difficulty and capability area.
Not every item is included in every collection — availability is confirmed during the sample review.
Where this data gets used.
Coding-agent post-training
Fine-tune coding agents against high-difficulty terminal and repository tasks.
Supervised fine-tuning
Use verified task and solution pairs as supervised training data.
RLVR & reinforcement learning
Verifiable-outcome tasks suited to RLVR and RL training pipelines.
Private model evaluation
Run confidential evaluations that avoid public benchmark contamination.
Capability-gap analysis
Identify where a model's coding-agent capabilities fall short.
Grader & reward-model research
Develop and validate graders and reward models against varied tasks.
Agent robustness testing
Probe agent behavior against difficult, adversarial-adjacent tasks.
Benchmark extension
Extend an existing benchmark's coverage with additional private tasks.
Failure-mode analysis
Study how and where agents fail across a broad task distribution.
Ways to acquire access.
Non-Exclusive Licence
Use the dataset while Deaimer retains the ability to license it to additional customers.
Exclusive Licence
Negotiate exclusive rights for a defined collection, capability area, language, format, market, or period.
Direct Acquisition
Acquire an agreed dataset package and associated rights under a negotiated asset-purchase or data-licensing agreement.
Final scope, permitted uses, exclusivity, warranties, delivery formats and intellectual-property rights are defined contractually.
From inquiry to delivered package.
Submit an inquiry
Tell us about your intended use and the collection you're interested in.
Execute an NDA
A mutual NDA is put in place before any detailed material is shared.
Review the catalog
Review a catalog and metadata summary describing available tasks.
Inspect sample tasks
Review a small set of controlled sample tasks from the collection.
Diligence & overlap checks
Agree on technical diligence and dataset-overlap checks with your team.
Select a structure
Choose a licensing or acquisition structure that fits your use case.
Receive the package
Receive the versioned delivery package under the agreed terms.
Inspect the data before discussing a larger transaction.
Request a structured catalog and a small set of controlled samples under your preferred confidentiality process.