The team cannot freely test alternatives, compare them under identical conditions, or calculate the full cost of an accepted result.
There is no room for free experimentation, no shared benchmark set, and no full cost for a result. Processes therefore evolve through intuition and personal habits rather than as a measurable engineering system.
The comparison should not be “bot or no bot,” but versions of the reviewer pipeline run on the same set of pull requests: measuring useful findings, noise, tokens, and human review time.