# Community questions-to-verification workflow

## Purpose

Turn repeated public-place questions into structured research work while keeping community discussion, evidence, candidate facts and published facts as separate epistemic layers.

## State flow

```text
public discussion (always unverified)
  -> privacy redaction and moderation
  -> group by place + topic
  -> research need
  -> attach qualifying evidence
     -> no evidence: needs_research
     -> one field report: supported_pending_review
     -> conflicting values: conflict_review
     -> current official source OR two independent matching field reports: verified_candidate
  -> maintainer review
  -> only an external authorized workflow may publish a durable place fact
```

The prototype deliberately ends at `candidate_requires_maintainer_approval`. It never mutates a production fact table.

## Entities

### Public discussion

Stores a pseudonymous author alias, sanitized body, place, time, detected topic, moderation status and `public_discussion_unverified`. The raw body is not copied into generated outputs after privacy processing.

### Research need

Groups visible questions by `placeId + topic`, retaining all question IDs and a representative sanitized question. A group count helps prioritize repeated community uncertainty without making popularity evidence of truth.

### Evidence

- `official_source`: HTTPS authority URL plus `checkedAt`; freshness window 365 days.
- `field_report`: dated observation plus pseudonymous independent `reporterId`; freshness window 90 days.
- `community_signal`: useful for prioritization only and categorically excluded from fact verification.

Every evidence item names a `claimKey` and typed `value`. Invalid provenance enters review. Stale evidence remains auditable but cannot promote a candidate.

### Fact candidate

Created only when a current official source supports the value or at least two independent fresh field reports agree. It carries evidence IDs, citations and the immutable state `candidate_requires_maintainer_approval`.

### Review queue

Contains invalid provenance, stale evidence, conflicts, single-report corroboration gaps, held content and privacy-redaction confirmation. Each item includes a reason and next action.

## Conflict policy

Different values for the same `claimKey` among qualifying evidence force `conflict_review`. This applies even when one item is an official source. The summary names the conflict, cites every item and prohibits publication until a reviewer resolves it with a newer authority source or documented field check.

## Citation-bound summary contract

An AI or deterministic summary may:

- restate only claims present in attached qualifying evidence;
- cite evidence IDs inline;
- say “limited evidence suggests” for a single field report;
- explicitly state unresolved conflict or absence of evidence.

It may not:

- infer a missing value from community agreement;
- conceal a conflicting source;
- convert an old observation into current status;
- publish a fact or remove maintainer approval;
- expose raw private contact information.

The included prototype uses a deterministic summarizer so its guarantees are testable. A future LLM adapter should validate its output against the same evidence-ID allowlist.

## Moderation and privacy safeguards

1. Email addresses and North American phone numbers are redacted before grouping.
2. High link density and repeated-character spam are held before public grouping.
3. Reporter identity is pseudonymous; no email, device ID or exact home location is required.
4. Field reports describe park conditions, not identifiable people.
5. Community signals cannot verify facts.
6. Audit digests detect unreviewed changes to discussion or evidence fixture sets.

## Upstream integration plan

1. Add `place_questions`, `research_needs`, `research_need_questions`, `evidence_items`, `fact_candidates`, and `verification_reviews` tables with row-level policies.
2. Run redaction/moderation before inserting a question into the public grouping job.
3. Recompute grouping deterministically on accepted questions; do not overwrite comment history.
4. Execute evidence evaluation in a server-side job with a pinned policy version.
5. Render separate “Community questions”, “Research status”, “Evidence”, and “Verified information” UI regions.
6. Require a maintainer decision with actor, timestamp and evidence snapshot before copying a candidate into place facts.
7. Add migration fixtures and the included behavioral tests to CI.

## Acceptance mapping

| Requirement | Evidence |
|---|---|
| Comments remain distinct from facts | Separate `publicDiscussion` and `factCandidates`; invariant test |
| Repeated questions can be grouped | Place/topic grouping with preserved question IDs; test and demo count |
| Official sources and dated reports support answers | Typed evidence validators and freshness rules |
| Conflicts are flagged | `conflict_review`, high-priority queue item and fixture |
| AI summaries cite evidence and uncertainty | Citation-bound deterministic summaries and tests |
| Moderation/privacy documented | Redaction, hold rules, pseudonymous reports and safeguards above |
| Tests or reproducible prototype | Dependency-free CLI, 13 tests, verifier, fixtures and browser demo |

## Known limitations

- Topic classification is intentionally rule-based and English-only.
- Semantic grouping is place/topic level; production should add reviewed embeddings or synonyms without losing deterministic fallbacks.
- Redaction patterns are a bounded baseline, not a complete personal-data detector.
- The prototype does not authenticate reporters, evaluate source authority beyond type/provenance, or perform production writes.
- Fixture claims demonstrate behavior and must not be imported as current place facts without a fresh maintainer review.
