# Aaron, J1 and J2: what conditional probabilities can tell us

8 September 2026. Exploratory sensitivity analysis, separate from the website draft.

## Conclusion for Mike's reading

The priestly evidence supplies a worthwhile route for investigating continuity from Aaron. Within the accepted biblical paternal relationships, one securely identified living male-line descendant of Aaron would also establish surviving male-line descent from Levi, Noah and Adam. An enormous demographic estimate is not needed for that inference.

The calculations below do not identify such a descendant. They show that conditional odds can be calculated, sometimes strongly favouring J1, but their strength depends on assigning the observed priestly concentration to Aaron rather than to another old priestly founder. The surviving data do not supply a calibrated probability for that assignment. This is a limitation on historical identification, not a rejection of the biblical relationships being explored.

A finding that an authentic descendant belongs within the ordinary Y tree would favour an ordinary-tree placement for his paternal ancestors. In Mike’s usage, “copied or matched” includes creating the matching sequence afresh. Ordinary-tree placement would support a match at the level of a genetic group; it would not establish that Adam’s complete chromosome exactly matched a particular man living at his creation.

## Evidence used and scope

The previously audited figure in Hammer and colleagues' 2009 study contains 215 Cohanim: 99 J1, 63 across all J2 categories, and 53 in other haplogroups. Source: [Hammer et al., 2009](https://link.springer.com/article/10.1007/s00439-009-0727-5). This is a published volunteer sample, not a census or a representative probability sample of all living Cohanim. Pooling all J2 is deliberately coarse; it does not make those men a single recent paternal family. J1 is also broader than any proposed Aaron lineage.

The companion [research report](REPORT.md) discusses sequence studies, cross-community families and historical priestly records. Those observations motivate taking continuity seriously. They are not converted here into invented independent evidence multipliers. The numerical calculation uses the 2009 counts alone, once. It does not claim to incorporate all historical and sequence evidence into a calibrated likelihood.

## First calculation: conditional J1 versus J2

Assume Aaron's sampled family belongs to J1 or J2, and that its survival produces additional representation in that category. Start with equal prior odds, 50:50. Allow uncertain background frequencies rather than assuming every modern carrier of a category descends from Aaron.

The likelihood is Dirichlet-multinomial. For counts n=(99,63,53), let sample-category probabilities q have prior distribution

q | H_k ~ Dirichlet(tau b + c e_k).

Here b is the background mix, tau controls confidence in that mix, and c assigns additional prior weight to the proposed founder's category. These are analyst-supplied scenario assumptions, not estimates from historical sources or premises supplied by Mike. In particular, c is not a measured number of descendants.

For a = tau b + c e_k, A = sum(a), and N = 215, the integrated likelihood is

L_k = N! / product(n_j!) × Gamma(A)/Gamma(A+N) × product[Gamma(a_j+n_j)/Gamma(a_j)].

The posterior for J1 is L_1/(L_1+L_2). No best-fitting parameter is selected after inspecting the result.

| Background J1 / J2 / other | tau | c | Posterior J1 | Posterior J2 |
|---|---:|---:|---:|---:|
| 25% / 25% / 50% | 20 | 20 | 99.953% | 0.047% |
| 20% / 30% / 50% | 20 | 20 | 99.998% | 0.002% |
| 40% / 20% / 40% | 20 | 20 | 90.266% | 9.734% |
| 40% / 20% / 40% | 100 | 20 | 24.624% | 75.376% |

Across the 27 combinations of those backgrounds and tau,c in {5,20,100}, J1 ranges from 24.6% to almost 100%. This range is not a confidence or credible interval. It demonstrates dependence on uncalibrated modelling choices. The calculations show how J1 can be favoured; they do not justify selecting a particular percentage as the historical odds.

An especially important check: in the first row, the combined founder model is only 0.687 times as likely to produce these counts as its background-only model. Strong preference for J1 over J2 within the founder model does not mean the counts strongly favour that founder model itself.

## Second calculation: is Aaron represented at all?

A useful historical distinction is S, Aaron's line is sampled; U, it survives but is unsampled; and X, it has no living male-line descendants. Noah could have surviving descendants under any of these through other lines.

The following count model is narrower than those historical hypotheses. It conditions on a sampled Aaron line being J1 or J2 and being the enriched family. It leaves out sampled Aaron lines in other haplogroups and sampled but rare Aaron lines. Its results must not be reported as probabilities for the unrestricted S, U and X.

Let L_F=(L_1+L_2)/2 and L_0 be the background-only likelihood. Under U or X, permit another priestly founder to produce the same count distribution with probability r:

L_U = L_X = r L_F + (1-r)L_0.

Use prior probabilities S=0.50, U=0.25, X=0.25, with r=0.50. These prior choices are illustrations, not historical estimates. Applying Bayes' rule to the same counts gives:

| Background J1 / J2 / other, tau=c=20 | S, restricted sampled-Aaron model | U | X |
|---|---:|---:|---:|
| 25% / 25% / 50% | 44.88% | 27.56% | 27.56% |
| 20% / 30% / 50% | 60.79% | 19.60% | 19.60% |
| 40% / 20% / 40% | 2.93% | 48.53% | 48.53% |

These differing answers are not evidence that the true probability is somewhere between them. They demonstrate why choosing a model without historical calibration would manufacture apparent precision. U and X remain in their prior ratio because this model gives them identical likelihoods. The sample cannot distinguish those histories here.

The full grid contains 324 joint scenarios. It uses the counts once; the conditional J1/J2 result is a decomposition of S, not additional evidence multiplied into S.

## What would justify high confidence?

Separately, for any specified historical evidence E, posterior odds equal prior odds times its likelihood ratio. Starting with a 50% prior on sampled Aaron continuity:

| Assumed likelihood ratio for the evidence | Posterior |
|---:|---:|
| 1 | 50.0% |
| 2 | 66.7% |
| 5 | 83.3% |
| 10 | 90.9% |
| 19 | 95.0% |

These likelihood ratios have not been measured. The script supplies coherent illustrative Bernoulli likelihoods, P(E|S)=0.8 and P(E|not S)=0.8/LR, to make explicit what each calculation assumes. This illustration is separate from the count analysis.

If an alternative old priestly founder can produce exactly the same observations, and that possibility receives prior weight r under no sampled Aaron continuity, then L_notS is at least r L_S. The likelihood ratio cannot exceed 1/r within that model. With r=0.5, even maximally favourable evidence of this kind cannot move a 50% prior above 66.7%. This is a mathematical example of the identity problem, not a measured bound on the real historical probability. Evidence that distinguishes Aaron from that alternative would change the comparison.

## What further work would improve the answer?

A stronger model needs branch-level sequence data and ascertainment information; independently justified background frequencies relevant to the histories being compared; and historical constraints on the formation, transmission and expansion of priestly families. Modern national frequencies alone are not ancient priestly background frequencies. A documented paternal connection or independently identified ancient priestly DNA could be particularly valuable, though priestly identity itself would still need assessment.

Historical records and cross-community distribution can constrain alternatives without furnishing a complete pedigree. They should be investigated for that purpose. But multiplying subjective judgements for correlated pieces of evidence would not resolve the uncertainty.

For the website, the warranted positive conclusion remains: priestly paternal families provide a promising route to exploring surviving descent within the biblical genealogy. J1 and J2 deserve investigation. We cannot yet attach defensible historical odds to their identification with Aaron or Noah, or infer that copying Adam's chromosome was necessary.

## Reproducibility

- `bayesian_sensitivity.py`: complete likelihoods, priors and calculations.
- `branch-sensitivity.csv`: 27 conditional comparisons.
- `joint-sensitivity.csv`: 324 joint comparisons.
- `continuity-sensitivity.csv`: 12 illustrative historical-evidence scenarios.
- `bayesian-results.json`: selected results and assumptions.

Checks: posterior probabilities sum to one; branch probabilities sum to one; when r=1 makes sampled-Aaron and alternative histories observationally identical, posterior sampled-Aaron probability equals its prior throughout the grid. The earlier frequency counterexample remains preserved separately. No original study or website page was altered by this analysis.
