
Nichebench: the harness
So, Where Have I Been? π§
It's been about 10 months (!) since the last article. Most of that time went into my work at Dropsolid as an AI Product Architect. I've participated in the design, building, and maintaining the AI infrastructure: inference, models, guardrails, tracing, vector stores, and their integrations with Drupal. Quite a few moving parts - plus, some participation in the Drupal AI community, which was a fun process but definitely took some time β³.
Meanwhile, we secured β¬2.5 million in financing at Dropsolid AI. There was a lot happening on that side, and moving NicheBench forward alongside it was harder than I expected π’.
The hardware was another limit. Progress with one GPU was difficult, and even after adding a second one, running a model through the new benchmark could turn into days of trial and error - i.e. I'd start a run midnight and would wake-up to a failed attempt π₯. And, keeping up with the AI news takes a ton of time as well: new models, new harnesses, and experiments to understand what they could actually do. Useful work, but it competed with both NicheBench and writing about it.
So, this is where things stand. The goal is still to experiment with fine-tuned models for Drupal coding. NicheBench now has a runtime mode, QuietBee has more hardware, and I have some results to share - along with a few runs that went nowhere.
First, A Thank You π
Dropsolid AI became the first sponsor of our Drupal-focused AI experiments and this series of articles. They chipped in toward QuietBee's π second GPU, which let me run better models and test them on NicheBench's runtime tasks. That was a practical help with the experiments - thank you, Dropsolid!

By the way, if you or your company would like to help fund this work, there are sponsorship opportunities too. The goal is to work toward fine-tuned Drupal coding models (initially) and document what works along the way - including what doesn't. I can't promise a finished model, but that's the direction I'm working toward β.
Reach out to me via consultation page or via my LinkedIn.
NicheBench v2: Can A Model Do The Drupal Work? π§ͺ
I've been calling this NicheBench v2, though the actual package version is 0.2.x (idk when I'll be ready to call it >=1.0), the runtime-capable release. The earlier quiz and code-generation evaluations are still there. The new part is testing what an agent actually builds inside a running Drupal environment. The harness is in the NicheBench repository.
A model can answer Drupal questions or generate a convincing implementation without delivering a working feature. I wanted to test that next step: give it a project and a task, let it work through OpenCode, then check what it actually did.
How Does The Runtime Test Work? πΊοΈ
The main question is simple: can a model work as a coding agent inside a known harness? NicheBench gives it a Drupal project and a feature task (a custom branch), then lets it work through OpenCode - we simulate an actual dev-like environment which gives the agent a ton of flexibility.
The project supports different task flavors. For Drupal runtime, five task branches provide pinned starting states; NicheBench puts the selected task into its own workspace, starts the Drupal environment with DDEV, and launches OpenCode in the cage. The agent reads and changes the project, runs commands, and checks its work. NicheBench collects the result, runs deterministic checks, and sends the artifacts to a separate judge (an LLM itself, with hints about what's supposed to look "right").

I've grouped OpenCode and DDEV together in the diagram because they work on the same task - the DDEV containers actually run alongside the cage, not inside it. Each task gets its own workspace and Drupal setup β we clean up afterwards and start fresh for the next one.
Here's what I gave the agent to build:
- π οΈ 001: Application wizard
- π οΈ 002: Company verification
- π οΈ 003: Saved searches and alerts
- π οΈ 004: Company-profile completeness gate
- π οΈ 005: Advanced job search
These are feature-sized tasks, not one-function exercises. Runtime evaluations ran sequentially on this host. I also tried to force the agents to touch some API layers that were recently introduced (so it plays on either their knowledge, or their ability to discover it from the Drupal's Core itself).
What Do The Scores Mean? βοΈ
For this task pack, the runtime hybrid score is 50% deterministic checks and 50% judge score. Deterministic checks test defined outcomes; the judge reviews the implementation and its artifacts. The separate acceptance gate fails if a critical check fails, even when the score is high.
The score combines deterministic checks and the judgeβs assessment - useful for comparing what the models built, but not proof that every feature works. If a run stops before scoring, I leave it unscored, rather than calling it a zero.
Early Runtime Results - And Their Limits π
These results answer the runtime question better than a quiz or code-generation score can. They are still a snapshot of particular models and runs, not a current leaderboard.
|
Model |
Runtime mean |
Quiz |
One-shot coding |
Coverage |
Qualification |
|---|---|---|---|---|---|
|
MiniMax M3 |
85.94% |
Not reported |
Not reported |
1 Γ 5 |
Provisional; harness defects affect original scores |
|
GPT-5.6-Terra |
76.85% |
95.17% |
96.33% |
1 trial across tasks 001-005 |
Scored reference; single trial |
|
MiniMax M2.7 |
69.60% |
77.24% |
66.06% |
2 trials Γ 5 tasks |
Main cohort |
|
Qwen3.6-35B-A3B |
58.80% |
80.34% |
62.63% |
2 Γ 5 |
Main cohort |
|
Agents-A1 |
57.51% |
81.38% |
35.25% |
2 Γ 5 |
Main cohort |
|
Ornith-35B |
57.39% |
83.10% |
49.50% |
2 Γ 5 |
One unscored attempt was counted as zero in the aggregate |
|
Qwen3.6-27B NVFP4/MTP |
52.79% |
88.62% |
62.75% |
2 Γ 5 |
Main cohort |
|
Qwen3.5-9B |
35.49% |
85.86% |
25.52% |
2 trials Γ 5 tasks |
Overall aggregate retained; runtime SD unavailable |
|
Gemma-4-26B IQ4NL KV16/16 |
29.44% |
88.97% |
57.04% |
2 trials Γ 5 tasks |
Overall aggregate retained; runtime SD unavailable |
|
Laguna-S-2.1 |
Unscored |
87.59% |
53.29% |
No completed evaluation |
No runtime mean |
|
Qwen3.8-27B, local |
Unscored |
88.97% |
83.92% |
No completed evaluation |
Diagnostic attempts did not produce a completed evaluation |
Quiz is answer accuracy; one-shot coding is the judge-score mean, not a pass rate. Both are side references from separate runs.
Sorted by runtime score, highest first. Laguna and local Qwen3.8 have no scored runtime result β I've left them out of the graphs.

The score blends deterministic checks and judge assessment. It is not the acceptance result.

Quick heads-up before comparing the bars π
- Main cohort: Two trials across five tasks. Gemma and Qwen3.5 now have a confirmed two-trial split, but their runtime SD is unavailable.
- GPT: One trial across all five tasks. Useful as a single-trial reference, not a two-trial comparison.
- M3: One trial across five tasks. Its original 85.94% is provisional because command-observation and brittle-check defects affected the scores.
An unscored task is not a zero-quality implementation. Some historical aggregates do count unscored attempts as zero, notably Ornith-35B; read that row with this qualification in mind.
The runtime speed is also an interesting story - but at times also misleading. Some models believed they were done - stopped early and got an average or bad score, others tried hard to polish their work, built a test-suite, etc and in result took way more hours per 1 task.
Here's what I learned from these experiments:
π Harness + Freedom results in bigger spreads, lots of noise. The model can take a sudden turn and randomly either decide that it's done, or resolve the same problem wildly different. But you still see smaller models performing worse than bigger models (or, better trained).
π Qwen 3.6 35B A3B model performed close to Minimax M2.7, I really hoped that Qwen3.8-27B (currently my favorite local model) - would show results closer to GPT5.6-Terra or Minimax M3, but I still failed to get any meaningful results in my tests(will update this article when I do).
π Spin-off models like Agents A1 or Ornith - didn't show dramatic changes here. In some cases the classic Qwen 3.6 35B performed even better.
π There's still room for improvement - again we're not aiming to build best generic coding model - we just want to make sure it performs BEST in a specialized given domain. These series of articles are still relevant and so is the effort, or so I hope π«‘
BTW, the previous NicheBench vs Drupal 10-11 article covered the quiz and one-shot code-generation results. In NicheBench + New models, I was already looking at tool use, long contexts, and multi-turn behavior before Drupal knowledge. This round puts that agentic part to the test - an OpenCode agent that can inspect and edit a project, run commands, and check its work in DDEV.
Runtime Failures π§ͺ
- Qwen3.8-27B, local: I spent roughly two weeks trying to get it through the benchmark. Token speed looked fine, but I never got a completed score. I tried different budgets and local setups; it seems like the model is afraid to fail and spends 10x time just to get everything perfectly right - which means that 1 task = more than 9 hour run.
- Qwen3.8-27B, GROQ: By my recollection, hosted inference cost about $15 in an hour or two, then I stopped without a completed score. I was worried that it will repeat the same pattern, spending 4-5 hours (and, thus $30-$50) per just 1 task.
- Provider / judge calls: Not every interrupted run points to bad code. Ornith task 003 could not be scored after a judge/network failure. M3's first task had interrupted streams and retries; the retained capture doesn't establish where the interruptions came from.
- Laguna-S-2.1: Runtime attempts ended through watchdog or hard timeouts, with no completed runtime result.
The model, harness, provider, and budget all play a part here π€·ββοΈ.
UPDATE: QuietBee Is Getting Its Third GPU π
The original QuietBee build article was published almost a year ago. Back then, I had the Threadripper base with 96 GB of DDR4, and buying three RTX 5060 Ti cards was still the plan. Now two RTX 5060 Ti 16 GB cards are running, I've bought the third, and RAM is up to 128 GB of DDR4.
Yes, 128 GB of DDR4 in this RAM-apocalypse. TrendForce has been reporting reduced DDR4 allocations as suppliers prioritize HBM and server memory. Building a budget lab hasn't got easier.
| Β |
Running Now |
Once The Third Card Arrives |
|---|---|---|
|
AI GPUs |
2 Γ RTX 5060 Ti 16 GB |
3 Γ RTX 5060 Ti 16 GB |
|
Total Physical VRAM |
32 GB across two cards |
48 GB across three cards |
|
System RAM |
128 GB DDR4 |
128 GB DDR4 |
The first 5060 Ti cost me around β¬440, second ~β¬580. The third cost β¬750, and that was the cheapest I could find. World went crazy π€―.
I also came across Kai's updated local-AI hardware comparison. Almost a year after my build article, he landed on 2 x 5060 Ti @ 16GB cards too. Nice to see someone else arrive at the same choice - though his comparison also makes the point that the software and how you split a model matter a lot.
More GPUs don't automatically become one big pool of memory. What we can run or fine-tune still depends on the model and how the software uses the cards. Having that said, I still have room for 4th GPU (my PCIe lanes allow it) - so, looking forward for anyone who wants to help out!

What Comes After The Benchmark? π
The actual next step is data collection and synthetic data generation. There is quite a lot of complexity in that process, so I want to start with some basics, fine-tune a model, and check whether we can make meaningful progress on NicheBench.
I don't need to build the entire data pipeline before testing that idea. For this first experiment, start with a small, checked set of useful examples, run a fine-tuning experiment, and compare the model against its original baseline under the same evaluation conditions. Then use what we learn to decide what data to collect or generate next.
The evaluation tasks also need to stay out of the training set. Teaching a model the benchmark answers and then reporting a higher score wouldn't tell us much about whether it got better at Drupal.
- π Collect Drupal examples and generate synthetic data.
- π Start with a small, checked training set.
- π§ͺ Fine-tune a candidate model.
- π Compare it with the original model on held-out tasks.
- π Inspect what changed and improve the data.
That's the next experiment. Can a focused training set make a model better at this niche, and does that improvement show up beyond the quiz? I don't know yet. But we now have a way to measure more than whether it can explain a Drupal API.
If you work on Drupal AI, I'd be interested in what kinds of coding tasks you'd want to see tested - leave a comment π.
Now that you reached the end π - another thank-you to Dropsolid, the first sponsor of this Drupal-focused series. Their support helped get QuietBee's second GPU into the rig.
- β If you'd like to support the experiments, there's Ko-fi. Donations go toward the GPU budget and more articles. I'm already eyeing a fourth 16 GB GPU β 64 GB of total VRAM across four cards, giving me more room to continue the research.
- π Company sponsorship is welcome too. If you'd like to support the Drupal-focused work, get in touch - we'll discuss the details privately.
- π οΈ Have a project in mind? Check out HumanFace Tech and book a meeting.
π‘ Inspired by this article?
If you found this article helpful and want to discuss your specific needs, I'd love to help! Whether you need personal guidance or are looking for professional services for your business, I'm here to assist.
Comments:
Feel free to ask any question / or share any suggestion!