Skip to content

Temporary planning review record. The report below is reproduced verbatim as returned by the independent read-only reviewer (Plan agent, Fable model, launched 3 October 2026 about 14:35 BST). Only this front matter and note were added. Line numbers refer to the package as it stood when the reviewer read it. Resolutions are in the round-2 resolution matrix.

SR review: systematic review methodology and product completeness

Path legend (absolute roots): - PKG = /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/ - PLN = /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/ - MAIN = /home/chris/workspace/syrf/main/ - "ledger" = PLN/review-form-owner-decisions-2026-10-02.md

Methodology sources used from my own knowledge are named inline: PRISMA 2020 (Page et al. 2021, BMJ 372:n71), PRISMA-S (Rethlefsen et al. 2021), Cochrane Handbook v6 (ch. 4 selection, ch. 5 data collection, ch. 6 effect measures, ch. 7 RoB), SYRCLE's RoB tool (Hooijmans et al. 2014), the CAMARADES quality checklist (Macleod et al. 2004), ARRIVE 2.0 (Percie du Sert et al. 2020), ASySD (Hair et al. 2023), Vesterinen et al. 2014 (meta-analysis of animal data), and standard inter-rater reliability literature (Cohen, Fleiss, Krippendorff, Gwet). Anything I could not check in the repository is marked UNVERIFIED.

1. Verdict

The plan is strong on the mechanics of evidence integrity (immutable revisions, no fabricated history, snapshot gold, as-of reproducibility, PRISMA collective authority) and, taken with the FEAT-011/FEAT-012 specifications, it will make SyRF's screening, reconciliation and PRISMA counting methodologically sound. It does not yet add up to a very high quality systematic review facility, because several things a reviewer must report under PRISMA 2020 and must do under Cochrane/SYRCLE practice are absent from every release: the inter-rater reliability basis is wrong or unspecified (statistics are "computed from canonical revisions" with no "initial independent decision" marker, and screening-level agreement is not in R5c's scope); full-text retrieval has no human workflow, so boxes 6/7/12/13 cannot be truthfully populated; there is no protocol/registration record and no search documentation beyond a name and a file; there are no risk-of-bias or reporting-quality templates (RoB is done today with ad-hoc annotation questions and per-outcome judgements have already been lost once); outcome extraction has no extraction-method provenance, unit vocabulary or domain validators and graph digitisation is dead; there is no analysis-ready export and no way to group multiple reports into one study for box 10; and there is no calibration/pilot-screening support. The five most important changes: (1) add an "initial independent decision" marker and screening-outcome exposure to C3 at F1, and make screening-level IRR (per profile, rotating raters) an explicit R5c deliverable; (2) give P1 a human full-text retrieval workflow and amend FEAT-011's terminal FullTextNotRetrieved lifecycle precedence; (3) add a protocol/registration/search-documentation record (PRISMA items 6, 7, 24; PRISMA-S) to P1/R3d and bind profile-version publication to a protocol-amendment log; (4) ship SYRCLE/CAMARADES/ARRIVE question templates with per-outcome items in R1a/R3d, and extraction-method provenance plus a units/error-type catalogue in O1; (5) add PRISMA arithmetic-consistency acceptance criteria, report-to-study linkage (amendment B) and a comparison-level export to R5b/P2/O1. None of these reopens a confirmed owner decision; most are additive and cheap relative to the engine work.

2. Findings

ID Severity Location Finding Evidence Recommended change
SR-01 Major PKG/integrated-plan.md:662-671 (R5c), PKG/acceptance-criteria.md:319-326, PKG/contracts.md:155-170 (C3), PKG/open-questions-and-assumptions.md:90 (Q-16) Inter-rater reliability has no defined observation basis. R5c computes agreement "from canonical revisions" and only separates gold-informed work (VS2). It never says which revision per reviewer is the observation. DP2 lets a reviewer correct an Exclude from history "while review is possible", extra votes (D1/D2 Allow) resolve conflicts, and a reviewer can infer the collective state from availability and lock messages; a correction made after a conflict became visible is not an independent observation but C3's three exposure states only cover accepted gold. Screening-level agreement (per profile, rotating reviewer pairs) is also not in R5c's scope: AG1–AG3 and AC-R5c-01 are annotation-answer rules (multi-select sets, N/A). Journals ask for screening IRR (Cochrane Handbook ch. 4 recommends reporting agreement; PRISMA 2020 item 8). ledger:542-553 (DP2), :45-46 (VS1/VS2), :803-814 (AG3); PKG/contracts.md:163 (exposure states cover "accepted snapshot" only); MAIN/docs/features/materialized-project-statistics/README.md:223 (kappa excluded from FEAT-024); PLN/review-lifecycle-gold-settings-proposal-2026-10-03.md:93-103 (missing-state contract, no basis rule) At F1, add to C3: (a) a per-reviewer per-context "initial independent submission" marker (first effective Complete/decision before any collective outcome for that study/profile was visible to that reviewer), and (b) a screening exposure record: at correction time, whether the collective outcome (Pending/Conflict/Included/Excluded) or any reconciler/adjudicator output was visible to the actor. In R5c: default IRR view = initial independent observations; current-decision view available and labelled; corrections after collective visibility labelled "informed (collective)". Add screening-decision agreement per profile (two-reviewer κ where pairs are fixed; pooled pairwise κ or Krippendorff's α where raters rotate; prevalence-adjusted variant for low-inclusion TA screening) to Q-16's method review and to AC-R5c. Add a fixture: a DP2 correction after conflict never changes the initial-observation κ.
SR-02 Major PKG/integrated-plan.md:689 (P1 row), PKG/acceptance-criteria.md:346 (AC-P1-03), PKG/domain-model.md:128,159 Full-text retrieval (PRISMA boxes 6/7/12/13) has no human workflow. P1 records "retrieval status with the PDF acquisition and processing programmes" and AC-P1-03 only asks that status changes are "recorded as events". Nothing says who sets Sought/Retrieved/NotRetrieved, with what reason (paywalled, author contacted, no response), or that a PDF in SyRF is neither necessary nor sufficient for "retrieved" (full text can be read outside SyRF; a PDF can be the wrong document). Without an explicit action, box 7 is always 0 and box 8 ≠ box 6. FEAT-011 also makes FullTextNotRetrieved a terminal lifecycle state that takes precedence over an Included outcome, contradicting the plan's "lifecycle = pipeline position" rule and amendment H's per-profile outcomes; the plan does not amend it. MAIN/docs/features/prisma-specification/study-lifecycle-and-source-taxonomy.md:241,249-250 (T4/T12/T13), :258 (terminal), :597-601 (precedence over Included), :461 (box 7 from lifecycle); MAIN/docs/features/prisma-specification/three-level-data-model.md:211-219 (separate FullTextStatus); PKG/prisma-amendments.md:139-153 (H is silent on retrieval); PRISMA 2020 item 16a/flow diagram boxes 6–8. Whether Study Management Processing already exposes a "not obtainable" state: UNVERIFIED. Add to P1: a study-level "Full text" action (Admin and, by stage grant, Reviewer): Sought (date), Retrieved (how: PDF in SyRF / read externally), Not retrieved (reason from a small controlled list plus free text; author-contact date), each an append-only StudyLifecycleEvent with actor. Default rule: TA collective Include sets Pending→Sought automatically; PDF attachment sets Retrieved as a suggestion a human can override. Add amendment M to FEAT-011: retrieval is fullTextStatus only, never a lifecycle state; delete T4/T12/T13 and the precedence rule at :597-601; box 7 derives from fullTextStatus = NotRetrieved for TA-included studies. Add AC-P1-09: a study TA-included, marked Not retrieved, never FT-screened appears in boxes 6 and 7 and not in box 8; acceptance includes the reason export.
SR-03 Major PKG/migration-adoption-rollback.md:38 (writer list), PKG/source-status-inventory.md:216, PKG/contracts.md:164 (import provenance), PKG/prisma-amendments.md:212-214 (K rule 5) Screening decisions imported from CSV columns and mapped to project members have no defined status in the canonical model. Today the upload wizard attributes imported decisions to SyRF investigators, so after R3a they would be indistinguishable from independent in-SyRF decisions: they would vote, count in IRR, and count as "screened in SyRF" for amendment K's no-double-count rule even though they are steps done outside SyRF. C3 lists "source system, legacy ID" only for migration provenance. MAIN/user-guide/studies/upload-search.md:38-54 (map columns to investigators and a stage); PKG/open-questions-and-assumptions.md:128 (E26 captures legacy screening writes); K rule 5 "external screening counts are allowed only for records not screened in SyRF at that phase" Define in C1/C3 at F1 an authority = Imported screening-decision kind with source system, import job, mapped investigator and independence = unknown. Rules: imported decisions count toward profile sufficiency only if the importing admin asserts they were independent (recorded choice); they are excluded from the default IRR view and shown separately; for PRISMA they count as "screened in SyRF (imported record)" so K's external counts for the same phase are refused for those records. Add a fixture to PRISMA fixture 1 or 3. Extend the same rule to FEAT-004 annotation import.
SR-04 Major PKG/integrated-plan.md:689-690 (P1/P2), PKG/domain-model.md:160, PKG/prisma-amendments.md:112-121 (F), PKG/open-questions-and-assumptions.md:66 (Q-26) No protocol, registration or search-documentation support. The product holds a "Protocol Url" and free criteria text; SystematicSearch holds name, description, file type and a living-search link, nothing else. PRISMA 2020 items 6–7 require per-source name, platform, date last searched, full strategy, limits/filters; item 24 requires registration details, protocol access and amendments with reasons; PRISMA-S adds deduplication method/counts and update searches. Amendment F says "protocol amendments append" but there is no protocol entity to amend, and Q-26 (profile re-publication) is not linked to an amendment record even though a criteria change mid-review is a protocol amendment. MAIN/user-guide/projects/settings.md:16-24; MAIN/user-guide/getting-started/glossary.md:19-20; MAIN/user-guide/projects/create-project.md:27; MAIN/src/libs/project-management/SyRF.ProjectManagement.Core/Model/SystematicSearchAggregate/SystematicSearch.cs:45-70,158-159; MAIN/user-guide/studies/upload-search.md:26-36; PLN/review-guided-setup-template-plan-2026-10-03.md:40 (wizard keeps "protocol link" only) P1: add to SystematicSearch (nullable, N-1 rule) searchDate, platform, strategyText/attached strategy file, limitsAndFilters, dateRange, updateOf (search round), alongside sourceType/sourceName; expose them in the upload wizard and the admin source-classification tool. R3d (or an earlier small release): a project "Protocol & registration" record: registry (PROSPERO, OSF, other; whether PROSPERO accepts the project's animal review scope is UNVERIFIED), registration ID/URL/date, protocol document link/version, and an append-only amendments log (date, what changed, reason, which profile/form version it corresponds to). Bind F5 so that publishing a profile version with changed eligibility rules requires an amendment entry (optional "why it changed" already exists for questions; make it required for profiles). Export all of this in the PRISMA manifest (R5b) and as a "methods summary" (see Improvements).
SR-05 Major PKG/integrated-plan.md:305-320 (R1a), :564-573 (R3d), PKG/acceptance-criteria.md:104-114, 252-260, PLN/review-guided-setup-template-plan-2026-10-03.md:73-104 No risk-of-bias or reporting-quality support anywhere in the plan. Today RoB is done with ad-hoc annotation questions (seed projects use a "Risk of Bias" category with randomisation/blinding items); the AI tool is mothballed; the user guide has no RoB page. The little-DOMS investigation documents per-outcome RoB judgements (SYRCLE items 6–8) being lost. SET1 asks for profile templates and an initial form, and the setup plan lists study descriptors/model/treatment/outcome categories but no RoB/quality template. A preclinical SR facility run by CAMARADES without SYRCLE RoB, the CAMARADES checklist and ARRIVE items as ready templates is incomplete, and PRISMA 2020 item 11 requires reporting the RoB tool and process. MAIN/docs/architecture/seed-data-quality-analysis.md:311; MAIN/docs/roadmap/product-features-roadmap.md:39; MAIN/docs/features/calculate-rob-authorization.md:66-67; PLN/little-doms-investigation.md:605-629; ledger:826-845 (TC1), :926-933 (SET1); grep of MAIN/user-guide for SYRCLE/risk of bias returned no relevant page Add to R1a's template catalogue (and R3d's guided route) curated, versioned question templates owned by CAMARADES methodologists: SYRCLE RoB (10 items, yes/no/unclear, with items 6–8 targeted at the Outcome Assessment entity so judgements are per outcome), the CAMARADES 10-item quality checklist (study level), and ARRIVE 2.0 Essential 10 (reporting quality). Template rules: items carry a semantic role (rob.domain, rob.judgement) so exports can produce a per-study per-domain matrix; per-outcome items must be answered per outcome entity (TC1 feature-required types). Add AC-R1a-08: importing the SYRCLE template into a project yields per-outcome items bound to Outcome Assessment; and an export fixture producing a domain × study matrix. Independence of RoB assessors (two assessors) is then covered by the ordinary target-2 form and reconciliation.
SR-06 Major PKG/contracts.md:508-526 (C14), PKG/integrated-plan.md:693 (O1), PKG/open-questions-and-assumptions.md:87 (Q-17), PKG/acceptance-criteria.md:369-373 Extraction provenance and data quality are under-specified for meta-analysis. (a) No "how the value was obtained" field (reported in text/table, estimated from a graph, calculated by reviewer, obtained from authors); Cochrane ch. 5 and Vesterinen et al. 2014 require flagging graph-derived data, and SyRF's graph digitiser is dead (graph2data absent from package.json), so "graph assignment" only links a region. (b) Units are a free string with no vocabulary, so a measure can carry "mm3" and "mm³" across cohorts; no validator for unit consistency within a measure (ODIR1 fixes direction per measure but not units). © Error type is SD/SEM/IQR only; CI (lower/upper), range, median with Q1/Q3, and "unknown/not reported" are missing from the legacy-compatible schema; "variation" is undefined. (d) No domain validators listed (SD ≥ 0, n integer > 0, events ≤ total, time monotone within a series, SEM/SD plausibility). (e) Sample size is duplicated (cohort n vs series n) with no rule for which is the analysis n. MAIN/user-guide/data-extraction.md:74-80,103-117; MAIN/user-guide/data-export/data-dictionary/quantitative.md:86-90,164-180; MAIN/docs/superpowers/plans/2026-08-10-af2-phase4-experiment-parity-hosts.md:22 (graph2data dead); PKG/source-status-inventory.md:94-102 (fact 9); PLN/outcome-data-migration-plan-proposal-2026-10-03.md:95-99; ledger:699-720 (OC1 "variation has not been defined") In E12/F-O: (1) add an observation-level extractionMethod role (reported / graph-estimated / calculated / author-supplied / unknown) and a series-level dataSource note; make graph-estimated the default when a graph region is linked. (2) Add a project unit vocabulary (controlled list with free-text fallback, SI-aware labels) and a validator "same measure, different unit" that warns and blocks binding at reconciliation. (3) Specify the legacy-compatible and event-count schema fields precisely before Q-17 closes: average {mean, median, other}; dispersion {SD, SEM, 95% CI lower/upper, IQR Q1/Q3, range min/max, none reported}; n at observation (with "same as cohort n" default and provenance); events/total for dichotomous; time with unit. (4) List domain validators in AC-O1-02. (5) Record a decision on graph digitisation: revive a digitiser as its own lane or declare out of scope; either way the provenance flag ships in O1.
SR-07 Major PKG/prisma-amendments.md:68-76 (B), PKG/open-questions-and-assumptions.md:92 (Q-23), PKG/integrated-plan.md:690 (P2) PRISMA box 10/16 "studies" versus "reports" cannot be derived because nothing groups multiple reports (papers) into one study. Amendment B fixes report identity but no release gives reviewers/admins a "these reports describe the same study" action; P2's Publication is bibliographic identity, not study linkage. In preclinical reviews multiple papers from one experiment are common, and Cochrane (ch. 4.6.2) requires collating reports of the same study. Today new_studies would count papers. MAIN/docs/features/prisma-specification/prisma-flow-diagram-mapping.md:56-70,174-179; PKG/contracts.md:463-465 (units stay distinct but no linkage operation) Add to P2 (or R5b with amendment B): a "Link reports to one study" admin/reconciler action producing a StudyLink group (append-only, with reason and provenance), surfaced in the study view and in exports; box 10/16 count groups as studies and members as reports; extraction forms can show linked reports' PDFs side by side (later). Add PRISMA fixture 9: two reports linked → 1 study / 2 reports in box 10. Ask Chris whether linking may also merge extraction (recommend: no; extraction stays per report with a group key).
SR-08 Major PKG/contracts.md:433-457 (C11), PKG/integrated-plan.md:650-660 (R5a), PKG/ui-coverage-comparison.md:105 Export is current/previous/as-of CSV of the same long/wide shapes; there is no analysis-ready output. The quantitative export is one row per timepoint per cohort with control flags on units; the analyst must reconstruct comparisons (which cohort is the control for which, shared-control handling, experiment grouping). The classification research withdrew "combined analysis" as out of scope, but an export shape that metafor's escalc() can consume (one row per comparison: m1i, sd1i, n1i, m2i, sd2i, n2i, timepoint, measure, direction, units, experiment, study, report) is export, not analysis. There is also no RIS/EndNote export of included/excluded sets for reference managers, and no metaAnalysisIncluded owner (SR-20). MAIN/user-guide/data-export/data-dictionary/quantitative.md:12-14; PLN/unified-annotation-classification-research.md:896-897; MAIN/docs/user-guide-drafts/FEAT-013-export-prisma.md:28-37,147-156 (planned exports list no comparison export) Add a lane release X1 "Analysis-ready exports" after O1 and R4c: (a) comparison export (gold by default, candidates optional) with explicit pairing rules derived from Experiment membership and control flags, shared-control rows flagged, direction and units carried, one row per comparison × timepoint; (b) a machine-readable codebook per export (question version, options, semantic roles); © RIS export of any study set (included, excluded with reason, duplicates). Keep SMD/NMD computation out of SyRF, but document the recipe in the user guide.
SR-09 Major PKG/acceptance-criteria.md:328-338 (AC-R5b), :358 (AC-P2-07), PKG/prisma-amendments.md:52-66 (A) PRISMA accuracy is tested box by box but not for internal arithmetic. Only box 3's count-consistency equation is an acceptance criterion. A trustworthy diagram needs: identification sums (fields 31–34); records_after_removal = records_screened + not screened (the remainder exists under early stop/batches and must be shown, since the PRISMA2020 R template assumes equality); records_screened = records_excluded + dbr_sought_reports per column; sought = not_retrieved + assessed + not yet assessed; assessed = excluded_with_reasons + included; reported external counts (K) reconciled against the same identities. MAIN/docs/features/prisma-specification/prisma-flow-diagram-mapping.md:204-213; PKG/prisma-amendments.md:63-66 (entering screening, batches); PKG/acceptance-criteria.md:334 (K totals) Add AC-R5b-08: every snapshot passes a published list of arithmetic identities per column, with an explicit "not yet screened / not yet assessed" remainder shown in the manifest and diagram footnote; mismatches block freezing unless an admin records an explanation (as K already does for its mismatch). Add fixture 8b: early-stopped batched review.
SR-10 Major PKG/open-questions-and-assumptions.md:91 (Q-22), PKG/prisma-amendments.md:101-110 (E), :139-153 (H), ledger:555-571 (DP3) The primary exclusion reason is not deterministic. DP3 derives a decision from eligibility answers and shows "triggering criteria"; H allows "one primary reason or several counted reasons"; nothing defines how the primary reason is chosen when several criteria fail. PRISMA box 9/15 needs one reason per report; Cochrane practice is "first failing criterion in a pre-specified hierarchy". Leaving it to each reviewer's free choice makes box 9 non-reproducible and inflates reason reconciliation. PKG/ui-coverage-comparison.md:77 (profile editor has presentation order); MAIN/docs/features/screening-annotations/README.md:119-146 (primary reason as a single-choice question) In C4/F5, make the profile's criteria order the reason hierarchy and add a profile rule "primary reason = first failing criterion in configured order" (default on; off allows reviewer choice). The derived decision then records the primary reason automatically (DP3 reasoning), reason reconciliation (DP5) compares the full failing set but PRISMA reports the primary. Add AC-R3b-09.
SR-11 Major No mention anywhere in PKG (grep for calibration/pilot screening/training returned nothing); PLN/review-eligibility-policy.md:579; PLN/qm-v2-context/qm-v2-implementation-history.md:406 (R031 "training rounds") No calibration or pilot-screening support. Cochrane ch. 4.6.3 recommends pilot-testing eligibility criteria and extraction forms on a sample with all reviewers before independent work; many teams do this by hand outside the tool. SyRF's own eligibility policy notes "training or verification may be legitimate reasons to permit more reviewers to decide" but the plan carries nothing. Calibration decisions must not count toward sufficiency or PRISMA, and must be excluded from IRR by default (or reported separately). As cited Add to R3c (or R3a as a step option): a "Calibration" step kind: a fixed sample (admin-chosen or random N) is offered to every reviewer regardless of target; decisions and answers are recorded with purpose = calibration, never vote, never qualify, never enter PRISMA; the step shows live agreement and per-criterion disagreement; an admin "promote calibration decisions to live" action exists but is off by default. Add C12 rule: calibration records are not pool-entry or screening events.
SR-12 Major (question) PKG/open-questions-and-assumptions.md:79 (Q-29), ledger:260-266 (RA5), :45 (VS1) The plan models only two extraction methodologies: independent dual extraction with reconciliation, and single-reviewer unreconciled (Q-29). The widely used third one, "one extracts, a second checks" (Cochrane ch. 5.5.1 accepts this with caveats; SYRCLE guidance similar), has no representation: VS1 shows reconciled answers, not another candidate's, and RA5 is independent. Teams will fake it by using target-1 plus informal review, producing gold with no provenance of the check. As cited Ask Chris (see Q-3). Recommended: a "Verification" step kind on target-1 forms: a second reviewer with the verify grant sees the single candidate's answers (exposure recorded; labelled informed), confirms or edits, and the result becomes an attributed gold snapshot with authority = Verified (distinct from Reconciled and from Q-29's accept-as-gold). IRR is not computed for verified forms; exports label "single extraction, verified".
SR-13 Minor (question) MAIN/src/libs/project-management/SyRF.ProjectManagement.Core/Model/StudyAggregate/Screening.cs:23-30; PLN/screening-specialised-annotation-research.md:529-554; no "maybe/unsure" anywhere in PKG or PLN (grep) Screening decisions are binary Include/Exclude. At title/abstract, most tools and many protocols allow "Unsure/Maybe" (treated as include-for-full-text), which lowers false exclusion and is what Cochrane recommends ("when in doubt, include at TA"). A reviewer who is unsure today must either Include (contaminating agreement with hedged includes) or skip indefinitely. MAIN/user-guide/stages/screening.md:30 ("click Next to skip") Ask Chris (Q-1). Recommended: a per-profile option "Allow Unsure" (default on for TA templates, off for FT): Unsure routes like Include for downstream availability, collective rules treat Unsure+Unsure as Pending-needs-third-vote (configurable), PRISMA counts it as not excluded at TA, IRR reports it as its own category (three-category κ) and as collapsed Include.
SR-14 Minor PKG/integrated-plan.md:679-680 vs PKG/prisma-amendments.md:193-194 Inconsistency: R5b says box 1 (updated reviews) "stays deferred", but amendment K's step types include "studies from a previous review version", which is exactly box 1's content. As cited; MAIN/docs/features/prisma-specification/prisma-flow-diagram-mapping.md:76-81,256-268 Either remove that step type from K, or (recommended) let K populate box 1 as reported counts and switch the template variant to "updated review" when such a record exists, with the data-model extension FEAT-011 already allows (PreviouslyIncluded). Update R5b, AC-R5b-07 and C12.
SR-15 Minor PKG/prisma-amendments.md:139-153 (H), PKG/domain-model.md:159 FEAT-011 rule 6 ("transition to Included only when all required profiles are Included") is kept, but T13 and the precedence edge case make FullTextNotRetrieved a terminal lifecycle state that overrides an Included outcome; the plan's Study changes list both lifecycleStatus and fullTextStatus without resolving the overlap. MAIN/docs/features/prisma-specification/study-lifecycle-and-source-taxonomy.md:241,249-250,258,597-601 Covered by SR-02's amendment M; list it in the amendments summary table and in PKG/decision-register.md §2 as superseded wording.
SR-16 Minor PKG/contracts.md:210-230 (publication command), PKG/open-questions-and-assumptions.md:63 (Q-34), PKG/acceptance-criteria.md:190 (AC-R2c-02) autoUpdate "carries compatible answers forward unchanged" on the admin's say-so. Methodologically, an answer given under old wording and counted under new wording is a reinterpretation; if the compatibility judgement is wrong the error is silent and propagates into agreement and synthesis. The plan records a compatibility relation but does not require a rationale, does not restrict autoUpdate to non-semantic changes, and exports do not show which qualification policy applied. ledger:122-139 (recovered baseline; "compatible" undefined); AG3 flags cross-version comparison but only in statistics Keep the recovered choices (not reopened) but add guards: (a) autoUpdate allowed only when the question version's compatibility class is "presentation-only" (wording/help/ordering) or an explicit per-option mapping (Q-34) that is one-to-one; many-to-one option mappings and any change to free-text questions force requireReanswer; (b) the publication record stores the admin's rationale; © exports (R2a previous-version and R5a) carry qualificationPolicy and answeredUnderVersion per answer.
SR-17 Minor PKG/contracts.md:433-457 (C11), PKG/integrated-plan.md:650-660 (R5a), ledger:915-923 (PR1), PLN/review-steps-prototype-handoff.md:266-271 (surplus) Under DP6/EW1, extraction evidence will exist for studies that are collectively Excluded or still Pending. PR1 protects the PRISMA counts, but the export contract never says that extraction datasets default to collectively Included studies, nor how surplus assessments are labelled. An analyst downloading "current answers" could include excluded studies in a meta-analysis. Whether today's annotation export already filters by screening outcome: UNVERIFIED. As cited Add to C11 (F6a) and AC-R5a: extraction exports default to studies whose required profiles are collectively Included; an explicit option includes others, with per-row collectiveOutcome, surplusAssessment and profileVersion columns. Same default for the X1 comparison export.
SR-18 Minor PKG/integrated-plan.md:594 (R4a: RE1 optional explanations), ledger:313-319 (RE1) RE1 (confirmed) makes an explanation optional even when the reconciler overrides every candidate. Methodological consequence: the gold record can contain unexplained overrides that no audit can distinguish from errors. Not reopening RE1, but the plan has no QC surface for it. As cited Add to the R5c agreement view (or R4a's pool page) a "reconciler overrides" count: gold ≠ all candidates with no explanation, per form/question, exportable; default the non-blocking reminder on. Add an export column goldDiffersFromAllCandidates.
SR-19 Minor PKG/prisma-amendments.md:226-268 (L), PKG/acceptance-criteria.md:352-360 Dedup QC is one-directional: auto-confirmed merges (AutoConfirmed tier) are applied before screening with no sampled human check and no reviewer-facing way to flag a duplicate noticed during screening (common in Covidence/Rayyan). ASySD's published specificity > 0.999 (Hair et al. 2023) still means some false merges at scale; those remove a record from screening silently. MAIN/docs/features/deduplication/service-specification.md:39-44,234 Add to P2: (a) a reviewer action "Flag as possible duplicate of…" that creates a DuplicateReviewItem; (b) an admin QC sample (configurable %) of AutoConfirmed groups shown in the review queue; © PRISMA manifest records the ASySD version, thresholds and the human-reviewed share.
SR-20 Minor PKG/integrated-plan.md:678 ("Box 17 needs an owner and UI"), PKG/acceptance-criteria.md:337 (AC-R5b-06), PLN/review-permission-matrix-proposal-2026-10-03.md:59-94 (no capability) metaAnalysisIncluded has an acceptance criterion but no release, UI or capability delivers it. It is a per-study analyst decision (with a reason: no usable data, outcome not reported, etc.), not a screening outcome. As cited Place it in X1 (SR-08) or R5b: a "Synthesis inclusion" study attribute (included / excluded with reason / not applicable) under a Record synthesis inclusion capability, exported and used by box 17; never derived from extraction completion (already a C12 rule).
SR-21 Minor PKG/integrated-plan.md:502-504 (R3a default profile "reproduces the project threshold"), PKG/acceptance-criteria.md:231 (templates), PLN/review-guided-setup-template-plan-2026-10-03.md:77-80 The default decision rules for new canonical projects are unspecified. The legacy maths (automated dual: inclusion ratio with 0.333 threshold, Include+Exclude insufficient, third vote decides) is reproduced for compatibility, but new TA/FT templates should state an explicit, methodologically defensible default: two independent screeners, unanimity required, conflict resolved by a blinded third screener or adjudication, FT exclusions require a primary reason. PLN/screening-specialised-annotation-research.md:167-181; MAIN/user-guide/stages/screening.md:72-77 Specify the template defaults in R3b/SET1 content (owner: CAMARADES methodologists): TA = dual independent, Unsure allowed (if SR-13 approved), conflict → third screener; FT = dual independent, reasons required (DP5 on), conflict → adjudication; single-screener mode labelled "not recommended for publication" (as the user guide already says). Record them as PROPOSAL thresholds for F5.
SR-22 Minor PKG/acceptance-criteria.md:352 (AC-P2-01) The parity criterion "identical AutoConfirmed groups" to the R package is likely unattainable (string normalisation, Unicode, locale differences) and measures the wrong thing; the methodological claim that matters is sensitivity/specificity against the labelled benchmark. Hair et al. 2023 report sensitivity 0.95–0.998 and specificity > 0.999 on labelled datasets (MAIN/docs/features/deduplication/service-specification.md:52-57) Replace with: on the published labelled benchmark datasets, the C# implementation's sensitivity and specificity are each within 0.5 percentage points of the R package's (PROPOSAL), and every divergent pair is listed for review; keep "99% pair agreement" as a secondary indicator.
SR-23 Minor (question) ledger:45 (VS1 candidate isolation), PKG/integrated-plan.md:526-527 (extra votes before R4p), PKG/contracts.md:359-389 (C9) Conflict resolution routes are extra independent vote or adjudication only. "Resolution by discussion between the two screeners" (Cochrane ch. 4.6) is absent and cannot be done inside SyRF without breaking VS1; teams with two screeners (the user guide's own FAQ case) have no in-tool path except a reconciler. MAIN/user-guide/stages/screening.md:73-80 Ask Chris (Q-2). Recommended: a per-profile resolution route "Discuss": after a conflict, both candidates may see each other's decision and reasons (exposure recorded, labelled informed), either may correct via DP2, agreement rules re-run; IRR uses initial observations (SR-01). Default off.
SR-24 Minor PKG/integrated-plan.md:679-680, PKG/acceptance-criteria.md:246 (R3c new arrivals), MAIN/docs/funding/index.md:66 (living search deferred) Living/updated reviews are only indirectly supported (new arrivals reopen stages; living search exists but is deferred/flagged off). There is no "search round"/review-version concept, so an update search cannot be reported separately (PRISMA-S item 13; PRISMA updated-review template), and K's box 1 (SR-14) has nowhere to hang. MAIN/docs/features/prisma-specification/three-level-data-model.md:398-400 (incremental dedup resolved) Minimal now (P1): searchRound on SystematicSearch and ExternalStepRecord; report snapshots filterable by round; R5b shows per-round identification. Full updated-review support stays deferred, but reserve it in C12 so box 1 can be computed from PreviouslyIncluded + round later.
SR-25 Note PKG/open-questions-and-assumptions.md:71 (Q-30 VS1 default off) Correct default for independence. One further blinding option is missing: hiding bibliographic identifiers (authors, journal, year) during screening, which some protocols require to reduce prestige bias; cheap as a per-profile display option. As cited Add as a profile presentation option in R3b (PROPOSAL, default off); record in exposure provenance that metadata was hidden.

3. Improvements (beyond defects)

  1. Methods summary generator (R5b). From the data the plan already captures (profiles, targets, decision rules, reconciliation method per form, DP5 setting, IRR, ASySD version, external steps, protocol record), emit a structured "Methods" block (JSON and prose) covering PRISMA 2020 items 5–11, 16, 24: number of reviewers, independence, how conflicts were resolved, agreement statistics, dedup method and counts, retrieval failures. This turns the audit trail into publication text and is where SyRF can differentiate.
  2. Near-miss excluded list (PRISMA item 16b). Add an export preset "FT-excluded studies with primary reason and reviewer/reconciler provenance" to R5b or R5a; the data exists once R3b/R4p ship. Add an AC.
  3. IRR over time (SR-01 follow-on). In R5c, show agreement per reviewer pair and per criterion over screening order (first 100, next 100…), so drift is visible; this is the calibration feedback loop for SR-11.
  4. Reason reconciliation economy. With SR-10's hierarchy, show in the R4p adjudication view the failing-criteria vectors of each candidate rather than only the primary reason, so DP5 reconciliation is a comparison of structured sets, not free text.
  5. Data-extraction QC view (O1/R4c). Per form: count of graph-estimated observations, unit-mismatch warnings, SD/SEM flips corrected at reconciliation, missing n. These are the fields reviewers get asked about by referees.
  6. Dedup transparency in the PRISMA manifest (P2/R5b). Record algorithm version, thresholds, auto vs reviewed share, reversals; PRISMA-S item 16 asks for this.
  7. Report linkage UI reuse (SR-07). The duplicate review queue's side-by-side pair view can be reused for "same study, different report" decisions with a different outcome (link, don't merge).
  8. Codebook export (SR-08b). Every export ships a machine-readable codebook (question identity, version, wording, options, semantic role, entity scope, requiredness); analysts need it to interpret versioned data (AG3) correctly.
  9. Guided setup content ownership. State in R3d that template content (profile criteria templates, RoB/quality templates, default decision rules) is authored and versioned by CAMARADES methodologists via the R1a template mechanism, not by engineers, with a review date per template.
  10. Default "two reviewers + reasons at FT" guard rails. Where a project uses single screening or target-1 extraction, the project overview and PRISMA manifest show a persistent "methods caveat" label; the user guide already warns, the product should too.

4. Questions for Chris (product decisions only)

# Question Recommendation
Q-1 Should screening profiles support an Unsure/Maybe decision at title/abstract (SR-13)? Yes, as a per-profile option, default on in the TA template; routes like Include; collective rule configurable; reported separately in IRR and collapsed into "not excluded" for PRISMA box 5.
Q-2 Should a discussion route (mutual visibility after conflict, correction via DP2) be an allowed conflict-resolution route (SR-23)? Yes, per profile, default off; exposure recorded; IRR uses initial decisions.
Q-3 Should SyRF support extract-and-verify (one extractor, one checker) as a first-class workflow producing verified gold (SR-12)? Yes, as a step kind on target-1 forms with Verified authority; labelled distinctly from reconciled gold in exports and in the methods summary.
Q-4 Should calibration/pilot rounds exist as a step kind whose records never vote, qualify or count for PRISMA (SR-11)? Yes, in R3c; promotion to live decisions only by an explicit admin action, default off.
Q-5 Should the facility hold a protocol/registration record and search documentation (SR-04), with profile re-publication tied to an amendment entry? Yes; search fields in P1, protocol record in R3d or an earlier small release; amendment entry required when eligibility rules change (F5).
Q-6 Are SYRCLE RoB, CAMARADES checklist and ARRIVE templates in scope for R1a/R3d, with CAMARADES owning the content (SR-05)? Yes; per-outcome items bound to Outcome Assessment; versioned templates.
Q-7 Full-text retrieval: who may mark Sought/Retrieved/Not retrieved, and is a PDF in SyRF ever automatically "Retrieved" (SR-02)? Admins and stage-granted reviewers; PDF attachment suggests Retrieved but a human confirms; Not retrieved requires a reason.
Q-8 Should report-to-study linkage (several papers, one study) be supported for box 10/16 (SR-07), and does linking affect extraction? Yes in P2 with amendment B; linking never merges extraction.
Q-9 Analysis-ready export (comparison-level, metafor-shaped) and RIS export: in scope as lane X1 (SR-08)? Yes; shape only, no effect-size computation inside SyRF; RIS export early because it is cheap.
Q-10 Graph digitisation: revive a digitiser as a separate lane, or declare out of scope and ship only the "estimated from graph" provenance flag (SR-06)? Ship the flag in O1 now; decide on a digitiser lane after O1 pilots report how often graphs are the only source.
Q-11 Should box 1 be populated from amendment K's "previous review version" step type now, switching the diagram variant (SR-14, SR-24)? Yes, as reported counts only; full updated-review support stays deferred.
Q-12 Agreement method (Q-16 refinement): approve the default basis "initial independent observations", screening IRR per profile, and pooled pairwise κ / Krippendorff's α for rotating raters, with percent agreement and prevalence always shown (SR-01)? Yes, after a statistician checks denominators; κ only once Q-16 closes, as the plan already says.
Q-13 Should the primary exclusion reason default to "first failing criterion in configured order" (SR-10)? Yes, default on, with a per-profile override allowing reviewer choice.

5. Coverage gaps

  • Full-text retrieval workflow and non-retrieval reasons (SR-02): no release, no UI, no actor; FEAT-011's lifecycle treatment conflicts with the plan and is not amended.
  • IRR observation basis and screening-level agreement (SR-01): C3 lacks initial-submission and collective-exposure markers; R5c scope is annotation-only.
  • Protocol, registration, amendments, search documentation (PRISMA items 6, 7, 24; PRISMA-S) (SR-04): absent from every release.
  • Risk of bias / reporting quality (SR-05): no templates, no per-outcome RoB guarantee, no RoB export shape; the only RoB machinery is a mothballed AI tool.
  • Extraction-method provenance, unit vocabulary, dispersion catalogue, domain validators, graph digitisation decision (SR-06).
  • Report-to-study linkage (SR-07) and analysis-ready/RIS exports, metaAnalysisIncluded owner (SR-08, SR-20).
  • PRISMA arithmetic consistency (SR-09) and the box 1 inconsistency between R5b and amendment K (SR-14).
  • Imported (CSV) screening decisions' authority, independence and PRISMA treatment (SR-03).
  • Calibration/pilot rounds (SR-11), verification workflow (SR-12), discussion resolution (SR-23), Unsure decision (SR-13): no representation; each needs an owner decision.
  • Default decision rules for new projects (SR-21) and primary-reason determinism (SR-10): unspecified before F5.
  • Export defaults by collective outcome and surplus labelling (SR-17); reconciler-override QC (SR-18); dedup QC sampling and reviewer duplicate flags (SR-19); living/update rounds (SR-24); bibliographic blinding option (SR-25).
  • Items I checked and found adequately covered (no finding): PRISMA collective authority (PR1, amendment A), per-profile outcomes with route provenance (H), reviewed-duplicate handling (D), no fabricated history and as-of reproducibility (C11, EX2), target-1 gold (Q-29), legacy reconciled answers (Q-35), self-reconciliation default (Q-36), N/A and version-compatibility rules (AG3), blinding of candidates (BL1/VS1), notification independence from obligations (C15), and the support-impersonation real-actor rule.

Critical Files for Implementation

  • /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/contracts.md
  • /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/integrated-plan.md
  • /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/prisma-amendments.md
  • /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/acceptance-criteria.md
  • /home/chris/workspace/syrf/main/docs/features/prisma-specification/study-lifecycle-and-source-taxonomy.md