The photo rating bench is a straightforward yet effective tool for image evaluation, specifically designed with the founder's preferences in mind. The setup is simple: put a batch of images in front of the founder using a local bench, where they can rate each image as no, mid-low, mid-high, or heart. This process is not for layout choices, which are handled separately by Hero Lab.
To initiate a session, create a bench folder inside the owning repository, such as <repo>/design-review/<topic>-<date>/. Store web-size copies of images (around 900px JPEG) there, avoiding full-size originals in Git. Alongside these images, write a frames.json file containing details like id and file, which are mandatory, and optional fields like reference and note.
The judging process involves running node scripts/scene-bench/qc-judge.mjs <bench> --model opus. This utilizes Claude's model to evaluate each frame against its reference, outputting results into qc.json. The file documents realism, AI tells, product truth, stare, and a verdict aligned with the founder's taste, which is calibrated and updated as needed.
To generate a review page, use node scripts/scene-bench/build-bench.mjs <bench> --title "...", which creates an index.html. Serve it locally with node scripts/scene-bench/serve.mjs <bench> 4321 and access it via http://localhost:4321/. All interactions, including votes and notes, are logged in <bench>/ratings.jsonl, ensuring no need for manual JSON handling.
Post-session, node scripts/scene-bench/agreement.mjs <bench> assesses agreement between the judge and the founder. Disagreements are key learning points; if they recur, these insights are incorporated into the judge's rubric and documented in CALIBRATION.md. This iterative learning is crucial, as the real challenge lies in how these adjustments perform in subsequent batches.
The process emphasizes local execution, following the founder's rule against hosted Claude Artifacts. Commit the essential files—frames.json, qc.json, ratings.jsonl, and index.html—to maintain a robust learning record. The judge's assessment serves as a filter; the founder's decision is final and always visible next to the judge's call.
By grounding this system in real interactions and feedback, the photo rating bench becomes a precise tool for aligning image selection with the founder's unique taste, ensuring that every decision is both informed and consistent.
Get weekly insights on AI architecture, pattern recognition, and building platforms without permission.
When dealing with documentation in software projects, ensuring that new or relocated docs are discoverable is as crucial as creating the content itself. I...
Read itIn the world of build systems, "cold-reopen" is an essential skill. It isn't about verifying something immediately after it's built; it's about returning to a...
Read itThe overnight build system's recent run highlighted the importance of handling routine tasks efficiently. The land operation in our system is not just about...
Read itHave thoughts on this post? I'd love to hear them! Join the conversation on X where we can discuss AI architecture, pattern recognition, and building platforms.
Discuss on XOr reach out directly at @TravisEric_