Early access: the directory is still filling out, and every rating here is a reported experience.

The journal

The job is to preserve signal

A review site has one asset, and it is not the reviews. It is the signal inside them: the true shape of what it is like to submit to a program, sharpened enough that the next researcher can act on it. Everything else (the grades, the rankings, the pages) is a way of reading that signal back out. Which means the failure mode that matters most is not a missing feature. It is quietly destroying signal you already had.

Almost everything shipped between August 5 and 11 was about not doing that. Some of it was new. A fair amount was us fixing our own mistakes. We are going to be plain about which was which, because a site that asks programs to be transparent has no standing to be otherwise.

One review was the wrong shape

Until this fortnight you could leave one review per program, and editing it overwrote what you had said before. That is a reasonable default for a product-review site. It is the wrong shape for this one, because the entire question here is how a program behaves over time, and a single overwriteable rating cannot hold time.

So a review is now a dated experience. The date is when it happened, not when you typed it up. You can log several with the same program, and they render as a timeline:

3★ 2023 → 5★ 2025 · improved

That line is the thing we could not show before. A program that was hostile three years ago and runs well now, and one that has always run well, used to collapse into the same single star rating. Now they read differently, because they are different, and a researcher deciding where to spend a month deserves to see which is which.

Letting people log many experiences opens an obvious hole: if every experience counted equally, one researcher with five reviews would drown out five researchers with one each. So the grading changed to close it. When a single researcher has several experiences with a program, their most recent counts for 80% and their earlier ones share the remaining 20%. Your voice still counts once. Stacking reviews cannot move a grade. And a program is graded on the number of distinct researchers who have reviewed it, not the number of reviews: the bar is people, not volume. Existing grades did not move when this shipped.

Deleting should not mean destroying

Two things used to destroy signal on a single click, and both are fixed.

Deleting a review used to erase it. Now it hides it, and a moderator can bring it back. Deleting your account was worse: it ran an instant, silent, final erase, and because reviews were tied to the account in the database, deleting the account took the reviews with it. One button, and a year of other people's reference material was gone.

Account deletion now deactivates (signed out, hidden from the site, but intact) and emails you a recovery link that works for 30 days. Change your mind, or discover someone else pressed the button, and one click brings everything back, reviews included. Full erasure still exists, because a genuine request to be forgotten has to be honourable. It is just no longer the accidental result of a mis-click. Every sensitive account event now lands in an audit trail, so "what happened to this account" has an answer instead of a silence.

Automatic checks can hold, but never reject

Reviews are checked before they publish, and we rebuilt what that check is allowed to do. A verified researcher's review can skip the manual queue by passing an automatic check, but the check has exactly two outcomes: publish, or hand it to a person. It can never reject anything on its own.

What it scores is completeness, and that word is doing real work. It is not length. A long review with an empty form does not pass; a tighter one that fills in the per-report detail, the ratings, and the outcomes does. We are measuring whether the review carries the information that makes it useful to someone else, not whether it hit a word count.

One rule sits above the score. A review that names an individual alongside an accusation is always held for a human, however complete it is. Deciding whether a specific named person did something wrong is not a call an automatic check should ever make, and we did not want a fast path that could make it by accident.

While we were in there we found and fixed a real bug: if you had already had a review approved, a new one was going straight to public with no check at all. The intent was that a trusted researcher skips the manual queue, not that they skip review entirely. Now every review lands pending and the automatic check picks it up within minutes. That was our mistake, and it is the kind worth naming.

Flagging the patterns that skew a grade, and only flagging

New detectors watch for three things: a burst of negative reviews on one program from a single account, the same text pasted across many programs, and a cluster of brand-new accounts all hitting one program at once. Each one raises a flag for a moderator. None of them unpublishes anything.

The last is deliberately, and only, a flag. Being a new account is not misconduct. Plenty of real researchers sign up and immediately review the program that brought them here. That is a normal thing to do, not a tell. More important: a detector that could auto-suppress a cluster of negative reviews is a weapon, and it points the wrong way. Any program that wanted its bad reviews gone would only have to argue that the researchers saying true things about it were a brigade. So the brigade detector does not get to suppress consensus. It gets to ask a person to look. Keeping grades trustworthy and protecting a program from criticism are not the same job, and we were careful not to build the second one by accident.

Email a provider will actually trust

This one was ours, and it was bad in a quiet way. Transactional mail (password resets, confirmations) was going out unauthenticated. It failed SPF and DKIM, so it landed in spam. Worse, providers disable the links in mail they cannot verify, so a reset email could arrive looking perfectly normal and still be impossible to use. If the email problem had already locked you out, account recovery could not reach you either. The two failures compounded.

Mail is now sent authenticated and signed as bugrater.com, so it reaches the inbox, and delivery is monitored: if it breaks again it raises an alert immediately, instead of failing silently the way it did this time.

Why say all this plainly

Several items here were features. Several were us repairing something that was plainly wrong. We are not going to sort them into a flattering pile, because the honesty is not a nicety on this site: it is the product. We ask programs to be candid about how they treat researchers. The only way to have any standing to ask that is to model it, including on the fortnights where the most useful thing we did was stop ourselves from destroying the signal we are here to keep.

← All posts