Proposed Reforms
What we don't actually know

The studies that would settle this

Most pages on this site end with a caveat: a range that's too wide, a follow-up period that's too short, a sample that's too small or too self-selected to lean on hard. Those caveats usually point at a specific, missing study — not a vague call for "more research." This page collects them in one place: what we don't know, what a study that actually answered it would need to look like, and who's positioned to run it.

What this page is not
  • Not a claim that the evidence is too weak to act on. Elsewhere on this site we argue the honest reading of current evidence already favours waiting for consent — this page is about sharpening that evidence further, not a precondition for the argument.
  • Not a wishlist of studies that would flatter one side. Several of these, done well, could easily cut against the position this site takes — that's the point of actually wanting them run.
What it is

A companion to the Proposed reforms page. Reform asks for procedural changes achievable now; this page asks a different question — what would we need to measure to stop arguing past each other on the specific points where the data genuinely runs out.

The gaps

Five studies that don't exist yet

1

A lifetime risk model that includes therapeutic circumcision

Every complication comparison on the Evidence page treats "circumcised" and "intact" as two fixed, separate populations with their own flat complication rates. That understates the real comparison, because a meaningful share of intact complications — severe phimosis, recurrent balanitis, paraphimosis — are ultimately resolved by circumcision, performed later, on a less favourable timeline, often under general anaesthetic instead of a local block.

A newborn circumcised at birth pays that surgical-risk cost once, upfront, for the whole population. An intact boy only pays it if he later develops one of the indications that actually leads to circumcision — and when he does, the procedure itself may carry a higher complication rate than the neonatal version being compared against. Neither side of that trade is currently modelled anywhere in the literature this site could find.

What it would take

Not a new trial — a modelling study, built from data that partly already exists: the proportion of intact boys who ever receive a therapeutic circumcision (by age and indication), and complication rates for circumcision performed at each later age band versus the neonatal baseline. Combined, those numbers would produce an actual lifetime expected-complication comparison, instead of two disconnected snapshot rates.

Who's positioned to run it: health economists or pediatric urology researchers with access to a large insurance-claims or national-registry dataset — the same kind of database already used for the complication studies cited on the Evidence page.

2

A sexual-function study that isn't small, self-reported, or confounded

As the Evidence page puts it plainly: nobody has won this outright. The measurement studies (fine-touch thresholds, quantitative sensory testing) suggest the removed tissue is genuinely the most sensitive part of the organ. The studies that try to translate that into lived sexual experience and satisfaction are, on both sides of the argument, small, self-reported, cross-sectional, or drawn from culturally confounded populations. The same bar belongs on the sexual-function literature around female genital cutting, which draws identical criticisms — small, self-reported, confounded — so that neither side of the broader genital-cutting debate gets held to a looser evidentiary standard than the other.

What it would take

A study that pairs objective quantitative sensory testing with validated, standardised sexual-satisfaction instruments, done prospectively rather than as a retrospective survey, in a population large and diverse enough to separate the effect of circumcision from the effect of the culture that made the decision.

Who's positioned to run it: sexual medicine researchers with access to a natural-experiment population — adult-circumcision cohorts, or countries with mixed intact/circumcised populations not sorted by religion, are the likeliest source of a design that isn't hopelessly confounded.

3

A real, prospective adverse-event registry that takes reports from adults

This one is already argued in full on the Proposed reforms page — it's proposal 7 there. The short version: the best complication numbers we have were reconstructed decades after the fact, from records nobody built for this purpose, because structurally nobody is required to track circumcision complications the way vaccines or blood transfusions already are. "If complications were common, we'd already know" only holds if someone is looking.

What it would take

A mandatory or opt-in complication-reporting registry, prospective rather than reconstructed, with a self-report pathway for adults — because some of the outcomes this debate is actually about (sensation, sexual function, psychological effects) surface at puberty or later, long after a newborn-period clinician has stopped watching and feedback is difficult to give.

Who's positioned to run it: professional societies, health regulators, or hospital systems — see Proposed reforms for the full case.

4

A longitudinal, cross-cultural study of psychological outcomes

The strongest evidence the Trauma page has for lasting psychological effects — Miani et al., 2020 — is a real, peer-reviewed finding, and the authors are honest about its limits: cross-sectional, self-reported, drawn entirely from a US population where circumcision status correlates with religion and region in ways that could themselves explain personality differences. It cannot establish that circumcision came first, in any causal sense.

What it would take

A prospective cohort followed from infancy into adulthood on standardised psychological measures, ideally spanning a population — like a country with a genuine intact/circumcised mix not sorted by religious practice — where the confound the current study can't rule out simply doesn't apply.

Who's positioned to run it: developmental psychologists running existing birth-cohort studies, who could add circumcision status and the relevant instruments to a cohort they're already tracking for other reasons.

5

A standard definition of the procedure itself — and metrics for when it's gone wrong

This sounds like it should already exist. It doesn't, not in any usable form. "Circumcision" is treated in the literature and in consent paperwork as a single, uniform procedure, but what actually happens on the table varies enormously: how much outer skin is removed, how much inner mucosa, whether the frenulum is preserved, reduced, or taken entirely, how much of the ridged band goes with it, how tight the resulting shaft-skin is left, how much is judged "too much" versus "not enough." None of that is standardized, and — this is the actual gap — there isn't an agreed set of outcome metrics that would let anyone say, after the fact, that a given result was appropriate versus botched.

That absence matters for the debate itself, not just for individual cases. The complication studies cited on the Evidence page mostly count things that are unambiguous — bleeding, infection, need for revision surgery — because those are the outcomes existing coding systems can capture. A result that falls short of a surgical emergency but still removed more than a functional amount, or left a degree of scarring or skin tightness that impairs later sexual function, has nowhere to be recorded. It isn't in the complication rate, because clinically nothing "went wrong." Whether an amount of tissue loss inside that normal range is itself a harm worth counting is exactly the question this site can't answer with a number — and neither, currently, can anyone else.

What it would take

A validated, standardized measurement and grading system for circumcision outcomes — analogous to grading scales already used elsewhere in surgery — that records how much outer skin and inner mucosa were removed relative to a defined baseline, frenulum status, residual shaft-skin mobility, and cosmetic outcome, collected prospectively across providers and techniques (clamp, Plastibell, freehand). Only once that exists does "botched" stop being a subjective judgment call and become a measurable deviation from a stated range — and only then can anyone meaningfully ask how often outcomes fall outside it.

Who's positioned to run it: pediatric urologists and surgical outcomes researchers, ideally the same professional societies that already publish technique guidelines — the gap is a measurement standard, not a lack of surgeons willing to be studied.

Why bother listing what we don't know

Uncertainty is a fact too

A site that only ever tells you what the evidence shows, and never what it can't yet show, is quietly overselling its own certainty. Every item above is a place where this site's own arguments rest on a caveat rather than a settled number — and naming the exact study that would remove the caveat is more honest than leaving it as a vague gesture at "more research is needed."

If any of these four studies got funded and run tomorrow, some of them might strengthen the case this site makes. Some might weaken it. That's what makes them worth wanting — not because we're confident how they'd come out, but because right now nobody actually knows.

Your turn

Know of a study already underway?

If one of these gaps is already being addressed — a registry being piloted, a cohort being followed up, a trial in progress — that's exactly the kind of correction this page should carry. The same goes for a fifth gap this list missed.

Add your contact or submission method here when the site goes live.