What a Co-Author's Pushback Is Actually For
Fifth entry in this series on the research process behind a voltage-screening pipeline. The first four posts covered the technical problems — a provenance chase, a circular threshold, a sample-size escalation, a step-function finding. This one is about where the pressure to find those problems actually came from: a co-author who never touched the code, and caught both real issues from a distance, purely by asking pointed methodology questions.
What the feedback actually said
It's worth quoting the substance rather than paraphrasing it, because the value was in how specific and checkable it was — not a vague "improve the rigor" note, but two concrete, falsifiable objections:
E3/E4 validate the loading–voltage relationship but generate no violation population, so they cannot yet validate screening recall under realistic profiles. Either we need a justified scaling/calibration that produces a meaningful violation population under E3/E4, or we should adjust the title and claims accordingly. In parallel, the 1.20 loading threshold must be explicitly separated from the 35-sample test if it was selected post hoc.
Both sentences turned out to be correct on inspection. The first is the provenance problem — real data producing zero violations, traced to a default table that was a documented overload scenario. The second is the circular threshold — a cutoff that turned out to equal the minimum of its own validation sample.
Why neither was visible from inside the work
Both problems have a specific character worth naming: they weren't hidden in complex code, and they weren't the kind of thing a test suite would catch. The circular threshold passed every check that was run on it — it just happened that the check itself was the thing that was broken. The provenance gap was similarly invisible from inside the pipeline: the default table worked, produced sensible correlations, produced a plausible violation count. Nothing about running it suggested there was a problem.
What both had in common was a kind of blindness that comes specifically from being close to the result. A number that confirms what you already expect the pipeline to do doesn't trigger the same scrutiny as one that surprises you — 239 violations under the default table looked like exactly the kind of thing a redispatch-screening pipeline should find, so there was no natural moment to ask "does this default table correspond to anything real?" until an outside question forced it.
The fork not taken
There was a real, faster alternative available once both issues surfaced: an "adopt and disclose" fix — cite the provenance gap and the circularity explicitly in a Limitations paragraph, keep the original headline numbers, and move on. That would have taken perhaps thirty minutes and technically satisfied the letter of the feedback: both issues would be disclosed.
The slower path — an original severity-sweep experiment, a full-population census, real out-of-sample regression testing — took considerably longer. It also produced a paper with an actual finding in it: the step function, where state-aware screening's value turns out to be conditional on how far conditions sit past the local violation onset. That's not a result "adopt and disclose" would ever have produced, because a disclosed limitation is still a limitation — it doesn't turn into new science just by being honestly labeled.
What it means to take review seriously
The easy failure mode when a co-author flags something like this is defensiveness — treating the objection as friction to route around rather than information to use. The harder, more useful response is to treat their skepticism as free signal: someone who hasn't seen the code and doesn't share your investment in the original number still found the exact place it was weakest. That's not an obstacle to the paper. It's the review process working correctly, and it's worth being grateful for rather than working around.
None of the four technical posts in this series would exist without that feedback. The provenance chase, the circularity check, the sample-size escalation, the severity sweep — all of it traces back to two sentences from someone who asked whether the numbers actually meant what they were claimed to mean, and didn't accept "it looks fine" as an answer until they'd been checked.
That's the methodology story — one that led directly to the paper's acceptance for presentation at IEEE ETECOM 2026 in Paris. The next entry is a shorter, different kind of finding — one that turned up while building supporting material, not while fixing a flaw.


