piker/plans/opencode/known-fqme-backfill-qualifi...

222 lines
8.4 KiB
Markdown
Raw Normal View History

# Known-FQME Backfill Qualification
## Purpose
Move `backfiller_deep_fixes` into the main checkout, reproduce the
historical backfill failures against one known-problem market, and turn
every confirmed failure into deterministic regression coverage before
expanding live qualification.
This is a correctness-debugging pass, not a broad provider smoke test.
One FQME and one timeframe remain fixed until their original failure is
understood, covered, repaired, and requalified.
## Current Repository State
The branch is currently checked out at:
`/home/goodboy/repos/piker/.claude/wkts/backfiller_deep_fixes`
Its head is `10e0ced7`. The linked worktree has only generated,
untracked commit-message metadata.
The main checkout is on `flake_update` at `c6ec3d41` and has staged and
untracked user work. Do not switch branches, stash, remove files, or
close the linked worktree until that state has been explicitly
preserved by the user.
## Phase 1: Transfer Branch Ownership Safely
1. Review the main checkout's staged and untracked state.
2. Preserve it using the user's chosen mechanism: finish and commit it,
move it to another branch/worktree, or explicitly create a stash that
includes the required untracked files.
3. Review the linked worktree's generated `.claude/skills/commit-msg/`
artifacts. Archive or discard them explicitly; they are not branch
source.
4. Remove the linked `backfiller_deep_fixes` worktree while keeping its
exact branch.
5. Switch the main checkout to `backfiller_deep_fixes`.
6. Verify all entry points import from the main checkout and resolve to
head `10e0ced7` or its descendant.
7. Run the 31-test deterministic analysis/backfill/storage set before
booting any daemon.
Gate: the main checkout owns `backfiller_deep_fixes`, tracked status is
clean, user work from `flake_update` is preserved, and deterministic
tests pass.
## Phase 2: Record The Known Failure
Capture these details for the selected case:
- provider and exact FQME;
- timeframe, initially `60` or `1` seconds;
- known-bad date/time range and timezone;
- original symptom: destructive truncation, duplicate seam, null rows,
out-of-order timestamps, hang, or another observed failure;
- command or chart workflow that exposed it;
- whether a known-bad parquet already exists;
- expected venue closure or continuous-session behavior;
- credentials, gateway, and network requirements;
- smallest reproducible provider response or persisted frame.
Do not start with multiple symbols, providers, or timeframes. Preserve a
checksum of any known-bad parquet before copying it into disposable
storage.
Gate: one case has an explicit expected result and enough source data to
distinguish a real provider/venue gap from a piker persistence defect.
## Phase 3: Build A Disposable Runtime
Use the main checkout's Python environment and a dedicated XDG config
root. The `pikerd` process and every client process must inherit the same
config root and import the main checkout.
The disposable directory contains only:
- a minimal `conf.toml`;
- credentials required for the selected provider;
- a dedicated `nativedb/` directory;
- an optional copy of the single known-bad parquet;
- captured logs and before/after audit output.
Ensure no older `pikerd` is listening on the selected registry address.
Verify `piker.__file__`, the branch, and the commit before starting the
daemon.
Gate: deleting the disposable directory cannot affect normal user
configuration, accounting data, or production history.
## Phase 4: Establish A Read-Only Baseline
Before backfill, record:
- parquet checksum and readability;
- earliest and latest timestamp;
- total and unique timestamp counts;
- duplicate count;
- zero and non-finite timestamp counts;
- non-positive timestamp deltas;
- null OHLCV counts;
- contiguous persisted index status;
- positive sample gaps with explicit endpoints.
Use a read-only audit helper or focused Python test. Do not automate
`piker store anal` or `piker store ldshm`; both can pause or mutate
attached data.
Gate: the original symptom is visible in saved evidence, or the plan
records why a fresh provider fetch is required to reproduce it.
## Phase 5: Reproduce One Lifecycle At A Time
Run in this order:
1. Load the known-bad persisted history without a provider write.
2. Start `pikerd` and request the chosen FQME/timeframe only.
3. Observe startup latest-frame query, pre-update TSDB snapshot, reverse
backfill, provider-delta commit, SHM publication, and notification.
4. Stop cleanly and rerun the read-only audit.
5. Restart and append to the same disposable history.
6. Compare earliest/latest timestamps, row counts, conflicts, and gaps
against the baseline.
Capture logs around every `update_ohlcv()` call, provider frame boundary,
SHM truncation, lock acquisition, temporary parquet, replacement, and
cache publication.
Gate: the run either preserves every unaffected row and normalizes the
known defect, or yields a bounded failure with enough data for a test.
## Phase 6: Convert Failures Into Tests
For each reproduced defect:
1. Save the smallest provider frames and persisted rows needed to
reproduce it. Prefer synthetic arrays; use a trimmed fixture only
when provider-specific boundaries matter.
2. Write a failing deterministic test before changing production code.
3. Assert exact timestamps, endpoint policy, persisted rows, and state
transition ordering rather than only asserting completion.
4. Keep actor-free tests in the existing focused modules where possible.
5. Add subprocess integration tests only for process death, advisory
locks, or atomic filesystem behavior that cannot be represented in
one Trio runtime.
6. Apply the smallest production fix and rerun both the focused test and
the complete deterministic set.
7. Rerun the exact manual FQME lifecycle that found the failure.
Potential deterministic modules:
- `tests/test_tsp_analysis.py` for normalization and classification;
- `tests/test_history_backfill.py` for provider/SHM/storage ordering;
- `tests/test_storage_nativedb.py` for merge and atomic durability;
- a new subprocess fault module for kill/restart and cross-process locks;
- provider-specific fixtures only when generic frames cannot reproduce
the behavior.
Gate: no production fix lands without a failing-before, passing-after
test or a written explanation for an unavoidable live-only condition.
## Phase 7: Fault Qualification
After the normal lifecycle is stable, exercise:
1. graceful Ctrl-C during provider query and storage update;
2. SIGTERM at deterministic query/write barriers;
3. SIGKILL after temp write, before replace, and after replace;
4. provider disconnect and recovery;
5. concurrent NativeDB writers for the same series;
6. restart after every interruption;
7. actual sampler backpressure and actor teardown;
8. MarketStore separately from NativeDB.
Use explicit barriers or failpoints rather than timing sleeps whenever
the failure position matters. Repeat the read-only audit after every
restart.
Gate: the final parquet is always old-valid or new-valid, unaffected
history remains present, indexes are contiguous, cache agrees with the
visible file, and shutdown remains bounded.
## Expansion Order
Only after the chosen FQME passes:
1. qualify its second timeframe;
2. qualify another continuous-session provider;
3. run Binance, Kraken, and Kucoin cases;
4. test IB with an available account and gateway;
5. add expected venue-closure classification;
6. widen to the remaining manual matrix.
## Commit Boundaries
Keep follow-up commits narrow:
1. captured fixture and failing regression;
2. minimal analysis/backfill/storage correction;
3. subprocess fault-injection harness;
4. provider-specific adapter correction;
5. qualification notes and reproducible audit output.
Every boundary must leave deterministic tests green. Do not combine
annotation performance, unrelated providers, or chart UX with the known
FQME debugging loop.
## Exit Criteria
The selected FQME/timeframe is complete when:
- the original bug reproduces from preserved evidence;
- a deterministic regression covers it where technically possible;
- fresh and append backfills preserve unaffected history;
- provider seam conflicts follow the documented policy;
- no zero, duplicate, non-finite, or non-positive timestamps persist;
- indexes remain contiguous;
- clean shutdown and restart preserve readable parquet;
- failure injection leaves an old-valid or new-valid file;
- remaining live-only risks are explicitly recorded.