REVIEW POLICY · Updated August 2026
How we test, score, and stay honest
Pipgen exists to answer one question without spin: which AI tools actually do the job? This page explains exactly how we run our benchmarks, how the scores are calculated, and why nothing on the site can be bought.
The core principle
Most “best of” lists rank tools the author never seriously used. They rerun a prompt until an image looks good, post the winner, and call it a review. We do the opposite, because a benchmark you can curate is just an opinion with numbers attached.
Every tool faces the same fixed set of prompts. We keep the first output, with no re-rolls. And we publish every image, wins and failures alike, so you can check our work rather than take our word.
How a benchmark works
Each lab on Pipgen has one fixed prompt set, chosen to probe the specific things that separate good tools from bad ones. For the AI image generator lab, that is nine prompts covering text rendering, photorealism, hands, product photography, anime and multi-part instruction-following.
- Identical inputs. Every tool gets the exact same prompts, pasted verbatim, on each tool’s default settings. We do not tune prompts per tool or use hidden system instructions.
- First output only. Whatever the tool returns first is what gets scored. No generating a batch and picking the best, no coaxing a better result out of a tool that fumbled.
- Full resolution. We grade at 100%, where a garbled sign or a sixth finger actually shows, not at a forgiving thumbnail size.
- Everything published. All outputs appear in the individual reviews, including the embarrassing ones. Nothing is quietly dropped.
How we score
Every image earns three marks out of ten, and the three are averaged into a per-image score. The image scores are then averaged into category scores, and the categories into an overall score.
The three axes
- Adherence — did the tool do what the prompt actually said? The right objects, the right words, the right layout.
- Coherence — does the image hold together physically? No melted fingers, no letters dissolving into glyphs, no shadows falling the wrong way.
- Aesthetic — setting the brief aside, does it actually look good?
What the numbers mean
- 9–10 — work you could ship as-is.
- 7–8 — a strong result with one visible flaw.
- 6 — something is clearly wrong on a close look, usually text or anatomy.
- Below 6 — the tool missed the brief outright.
Who runs the testing
Pipgen’s testing is run by Rajat Chauhan. Every result in the AI image generation lab is his own hands-on work: ten tools, the same nine prompts, the first output kept from each, and all ninety images scored against the rubric above. He studied at the University of Nottingham and works in AI services. You can reach him on LinkedIn, or email hello@pipgen.com.
Independence and how we make money
We sell none of the tools we rank, and no company can pay for a placement, a higher score, or a more favourable verdict. There are no sponsored reviews and no paid inclusions.
Outbound links go to each tool’s official site. We do not currently use affiliate links, and nothing on this site is a paid or sponsored placement. If that ever changes we will say so here and label the links, and it still will not influence the scores or the ranking: a tool’s position is decided entirely by its benchmark results.
Keeping reviews current
AI tools ship new model versions constantly, and a tool that was middling last quarter can jump after an update. We re-test as new versions are released, and each review notes the version tested and the date. When a re-test changes a score, the leaderboard and reviews are updated to reflect the most recent run.
Corrections
We aim to get the facts right, but if you spot a genuine error, a misquoted price, an outdated tier, a tool version we missed, we want to fix it. Corrections are made promptly and, where they materially change a conclusion, noted on the affected page. Pricing in particular changes often; we verify prices on the dates shown, and they may have moved since.