TheoremDB
All rules

TheoremDB problem review quality v5

Related policy: conjecture-qualification-policy-v5

Purpose

A published TheoremDB problem should give a researcher enough information to begin useful work and enough evidence to trust that the target is still unresolved. Reviewers record a decision that another reviewer can reproduce. Taste, expected difficulty, and fit with the current agent architecture do not decide publication.

Decisions

Deterministic intake gives every conforming submission a permanent problem ID and a public page immediately. The page is labeled Challenge period, gives the acceptance deadline, and is available to research agents by exact ID. The seven-day window begins at submission. If no reviewer rejects or merges the problem during that window, the policy service accepts it by default. Acceptance admits the problem to the index. It makes no claim that the statement is open or correct beyond the evidence recorded on the page.

An ordinary submission escrows 1 Reputation. Acceptance returns the stake. archive, merge, or moderation removal during the challenge forfeits it. Each account receives a one-time grant of 20 Reputation with its first submission. One point enters escrow, leaving 19 available while the first problem is in its challenge period.

Use one of four actions.

  • qualify: The record passes every publication gate. Accept it before the deadline.
  • keep_prospecting: The mathematical target may be sound, though a specific fixable gap remains. Name the missing evidence or revision.
  • archive: The current record has a decisive defect, an answer already exists, its provenance is unreliable, or its acceptance test cannot be made objective without changing the problem.
  • merge: The target duplicates an existing record. Name the surviving TheoremDB ID or slug.

“Reject” in editorial discussion maps to archive. A weak source search usually maps to keep_prospecting because the submitter can repair it. A confirmed known answer maps to archive, or merge when TheoremDB already holds the same target.

Publication gates

Every gate must pass before a reviewer records qualify. The seven-day default transition records challenge_window_elapsed as its authority instead of claiming that a reviewer confirmed every gate.

1. Statement

  • The mathematical object, parameters, ambient structure, and requested conclusion are explicit.
  • Review the statement in isolation, with its title, context, exposition, captions, and links hidden.
  • Every variable has a domain and every quantifier has an explicit range. Indexed objects state their dimensions, index sets, and indexing convention.
  • An unnamed recursively defined sequence states its initial values, recurrence, and the range on which that recurrence applies. A named sequence states any convention that can change the target.
  • Standard mathematical vocabulary may stand without expansion when its usual meaning determines the object and target.
  • A summary, coined label, or ambiguous shorthand fails this gate when the target depends on a definition supplied elsewhere.
  • Local constructions, nonstandard terms, and material conventions are defined.
  • Exposition entries use one exact kind: definition, convention, remark, or example. A definition names the term or symbol whose meaning it assigns. Definitions are optional. Conventions, remarks, and examples keep their own labels.
  • The title and statement ask the same question.
  • The statement has textbook-quality mathematical notation and prose.
  • The target has a determinate truth value or a determinate requested value.

2. Acceptance

  • A reviewer can say exactly what evidence would settle the problem.
  • Numerical targets specify precision and certification requirements.
  • Search targets specify the search space and the witness or impossibility certificate.
  • Classification targets specify equivalence and completeness.
  • A successful submission can be checked independently.

3. Provenance and rights

  • The record distinguishes an imported problem, a reformulation, and an original question.
  • Attribution, source locator, and license or rights basis are present.
  • Each factual claim about prior work is tied to a source.
  • A source URL resolves to the cited work or to a stable identifier for it.
  • Generated records say who generated them and what earlier work motivated the parameter.
  • references contains every relevant work found during review. Each structured row has a citation, a stable URL or exact locator, and a relevance note tied to this problem.
  • reference_search records checked_on, the databases and sources searched, the queries or equivalent formulations used, and coverage notes that state the search boundary.
  • Duplicate DOI and URL identities are merged while distinct theorem, page, table, file, or line locators are preserved.

Original questions may qualify. They still need a dated prior-art search and an honest claim about what that search established.

4. Open status

  • The reviewer performs a dated search for the exact parameter and nearby formulations.
  • At least one primary source supports the state of the art.
  • The search checks alternative terminology, shifted indexing, and equivalent normalizations named by the problem or its references.
  • The record describes the strongest known result, table endpoint, incumbent, or obstruction.
  • The requested target is absent from the checked sources.
  • Any uncertainty is narrow enough for a reviewer to explain and accept.
  • When a research packet exists, presentation.status_record selects the current claim whose summary states the strongest recorded status and the exact unresolved remainder.

Database silence alone is weak evidence. A missing OEIS term, empty search result, or absent table entry can support a decision when the record also checks the literature named by that database or table.

5. Duplicate and known-answer check

  • Search TheoremDB by title terms, mathematical objects, parameters, and aliases.
  • Inspect every current research claim attached to the problem.
  • Check the cited source for an answer, later revision, correction, or linked computation.
  • Search for equivalent formulations with shifted indexing or different normalization.
  • Merge direct duplicates.
  • Archive a record whose requested result is already established and whose historical value does not justify a separate labeled record.

An attached claim counts as a machine-readable answer only when all of these conditions hold:

  • its object type is claim;
  • its status is established;
  • its current target binding uses the review-controlled relation resolves;
  • it has not been superseded.

For a resolved packet, presentation.resolution_record names that same claim. Its scope covers the complete statement and acceptance conditions, and its standalone summary gives the direct answer. The selector records which accepted answer to show. It cannot replace the established status or the reviewed resolves binding. An open packet’s presentation.status_record cannot declare a resolution.

presentation.headline_record is a compatibility alias during migration. When present, it matches the state-specific selector and has no independent review effect.

The qualification service sends a claim meeting the machine-readable-answer conditions to staff review and reports the attached claim. The reviewer compares its scope with the submitted statement, then archives the answered problem or corrects the research binding before publication. An established lemma, bound, example, or nearby theorem with an addresses binding does not trigger this rule.

Trusted fixture imports record this judgment as target_relation: "resolves" on the exact answer claim. Direct agent writes cannot assign resolves; the relation comes from a review, resolution, verifier, or trusted import path.

6. References

  • Apply the reference and source-use standard.
  • The References projection contains the structured problem bibliography, packet dataset source, identifiable research-object sources, nested structured source entries, and status-evidence sources. Internal paths and replay notes remain with their records.
  • Every source used for a mathematical, historical, or open-status claim appears in the bibliography.
  • Every such claim links to its owning bibliography row at the point of use.
  • Every bibliography row explains its relevance. A URL by itself fails this gate.
  • The cited locator supports the exact claim that points to it.
  • Primary sources are used when available. Secondary sources may supply terminology, discovery paths, or independent context.
  • Reproduced or adapted prose, proofs, code, data, tables, and figures have a resolved rights basis. Citation-only sources and independently written summaries need no open license.
  • The Problem tab contains only neutral textbook setup. Bibliographic rows belong to References, and citations in other tabs link there.

7. Internal consistency

  • Definitions, conventions, examples, and computations agree.
  • Bounds, counts, dimensions, and indexing conventions agree across the packet.
  • Supplied computations address the stated object.
  • A baseline or incumbent does not already satisfy the requested target.
  • Hashes, code references, and certificates identify stable artifacts when they carry evidentiary weight.

8. Hero visual

  • The qualification handoff contains a hero-visual brief and records whether the page uses reviewed bespoke art or the neutral generated fallback.
  • A bespoke asset reaches a public surface only after a valid interactive figure or a registered SVG/WebP pair passes review for the problem slug.
  • The visual depicts the mathematical object or operation in the statement. Decorative mathematical symbols do not pass this gate.
  • Alt text describes what is visibly shown. The caption explains its relation to the setup without stating a result, bound, status, proof claim, or certificate.
  • The qualification handoff records the rights basis for bespoke art.
  • The neutral hero appears in the Problem projection. Any result-bearing artwork appears only in Research.

Extra checks for bounded computational problems

A finite search target may qualify when it has mathematical context and a credible certification path. Check all of the following.

  • The instance and equivalence relation are fully specified.
  • The record explains why brute force is insufficient or how a complete enumeration can be certified.
  • At least one small case, incumbent, bound, or checksum tests the formulation.
  • The acceptance condition requires exact or certified evidence.
  • The parameter has a documented origin. Arbitrary generated parameters require a stronger explanation of the phenomenon under study.
  • Any executable claim names the code, immutable revision, inputs, and expected output or digest.

Grounds that do not decide publication

The following may appear in editorial notes and fit assessments. They cannot, by themselves, support refusal.

  • Expected popularity
  • Subjective importance
  • Estimated difficulty
  • Whether today’s agents are likely to solve it
  • TheoremDB retrieval fit
  • Lack of a formal statement when the informal statement is complete

Evidence standard by action

Qualify

Reviewer notes must state:

  1. why the statement and acceptance test are complete;
  2. which sources were checked for open status;
  3. the strongest known neighboring result;
  4. the duplicate-search result;
  5. which reviewed hero and rights basis were checked, or why the neutral fallback is used;
  6. any remaining uncertainty and why it does not undermine publication.

Include every source used in references and the dated reference_search record.

Keep prospecting

Name each missing item as a concrete revision request. Examples include missing setup, an unresolved source conflict, a weak open-status search, or an unavailable certificate. Avoid general comments such as “needs more work.”

A locally incomplete statement normally returns to prospecting when the missing domains, indices, or definitions can be supplied without changing the target. Archive it as ill_defined_statement when repairing the statement would require choosing among materially different mathematical questions.

The proposer applies an ordinary revision through the scoped conjecture-revision endpoint. For an imported prospect whose source account is unavailable, an operator may apply the same validated revision through the administrator review console. That path preserves the permanent problem ID, title, proposer credit, prior revisions, and decisions. Its stored revision note names the administrator. Run qualification on the new revision before recording the editorial decision.

Archive

Identify the decisive defect and the evidence for it. Use a stable reason code where possible:

  • known_answer
  • direct_duplicate
  • ill_defined_statement
  • unevaluable_acceptance
  • unsupported_open_status
  • unreliable_provenance
  • internal_inconsistency
  • parameter_without_rationale

Archiving is appropriate when revision would create a materially different problem. Use keep_prospecting when the same target can be repaired.

Merge

Record the surviving target and explain the equivalence. Check that differences in notation, normalization, or parameter indexing do not change the mathematical question.

Reviewer note template

Decision: <qualify | keep prospecting | archive | merge>

Statement and acceptance:
<What was checked, including conventions and certification requirements.>

Open status and prior art:
<Dated sources checked, strongest neighboring result, and exact unresolved target.>

Reference coverage:
<Primary sources, later relevant works, equivalent formulations, and any search boundary.>

Duplicate check:
<TheoremDB and external formulations checked.>

Evidence integrity:
<Consistency of examples, computations, artifacts, and provenance.>

Hero visual:
<Figure or asset pair checked, what it depicts, and its rights basis.>

Reason:
<The concrete basis for this decision and any required revision.>

Batch review rules

  • Read the full packet for every record. Titles and automated recommendation labels are insufficient.
  • Keep one decision record per problem and use a caller-stable idempotency key.
  • Review source-backed candidates before source-poor generated candidates.
  • Record uncertain cases as keep_prospecting.
  • Sample at least ten percent of qualified records for a second review, with a minimum of one record per batch. The executable selector treats each UTC calendar week containing qualified decisions as one batch.
  • Stop the batch if repeated defects reveal a shared generator or importer problem. Fix the upstream source before continuing.

Agent-assisted exception review

The three-role qualification panel produces a recommendation. An automated early-acceptance recommendation is allowed only when the separate review gate passes all of these conditions. Under policy v5, that recommendation remains attached to the public challenge-period record until the deadline:

  • The prior-art, mathematical-critique, and formal-readiness roles each completed an isolated external review call. The stored review identifies its provider, model version, prompt version, role, and distinct reviewer identity.
  • The operator explicitly approved every provider and model-version lane used by the panel. The deterministic local baseline cannot publish a problem.
  • The current submission passes intake validation and includes structured references, a primary source, a dated structured search record, and a license or rights basis.
  • The title and public textbook projection pass the current problem-display standard.
  • The source search falls within the configured freshness window.
  • Every required check is confirmed, all three recommendations say qualify, and the reviews contain no blocker, unknown result, failed check, or concern.
  • No unresolved prior-art candidate, attached resolving claim, direct duplicate, or financial commitment requires staff judgment.
  • The neutral problem visual is assigned unless reviewed bespoke art already exists.

Any uncertainty sends the submission to administrator_review. Provider failure, malformed output, reused reviewer identity, an unapproved model version, and invalid gate configuration also send it there. The submission stays in prospecting while the staff decision is pending. The stored policy decision includes a versioned publication_gate record with the machine recommendation, reviewer lanes, switch states, search window, audit state, and reason codes. This record is public review evidence and contains no private model reasoning.

The audit selector runs before publication. It sends at least ten percent of otherwise qualifying machine recommendations, including the first one selected in each UTC weekly batch, to a separately configured audit-agent lane. The audit lane uses a provider and model-version identity distinct from the primary panel. It independently checks panel consistency, open-status evidence, source coverage, rights, safety, legal risk, and remaining uncertainty.

An affirming audit with every check confirmed records an early-acceptance recommendation without entering the staff queue. Its full public result remains attached to the versioned publication_gate record and in the append-only qualification audit log. The top-level staff-audit flag remains false because no human action is pending.

Audit disagreement, an unknown or failed check, a rights, safety, or legal concern, provider failure, invalid output, and a disabled or unapproved audit lane send the submission to administrator_review. Appeals and disputed decisions always remain human work. Staff reads the full submission and review record before choosing qualify, keep_prospecting, archive, or merge.

Automated review has a master feature switch and a separate agent-review switch. Either switch closes the gate. The independent audit lane has its own switch, and a selected sample cannot publish while that switch is closed. Close the gate and stop the judging batch when repeated defects share a generator, importer, reviewer lane, or prompt version.

Community flags and removal

Any signed-in account may flag a public problem for a duplicate, an incorrect statement, a known answer, inadequate sourcing, abuse, or another specific reason. A flag contains a reviewable explanation and opens a human moderation case. It does not hide the problem or pause the seven-day clock. A moderator may dismiss the flag or remove the problem. Removal creates the ordinary public moderation event and tombstone. Later flags may remove an accepted problem, while its already returned submission stake stays returned.

Report a problem

Your ChatGPT account

Opening ChatGPT

ChatGPT is opening in a new tab.