diff --git a/notes_to_self/claudy/gapsannotator_sesh.md b/notes_to_self/claudy/gapsannotator_sesh.md new file mode 100644 index 00000000..40c2471a --- /dev/null +++ b/notes_to_self/claudy/gapsannotator_sesh.md @@ -0,0 +1,122 @@ + + +● aight so perf analysis time, then i'll let you sleep for real: + + Current State (Final Benchmark) + + Server side: CRUSHED IT 🎉 + - 7.6ms total for 1285 gaps + - time_to_row creation: 1.3ms (was 98ms! 75x speedup from field extraction) + - Everything else sub-2ms + - Server is basically not the bottleneck anymore + + Client side: NEW BOTTLENECK IDENTIFIED + - 444.9ms total for 1285 gaps + - Building annotation specs: 324.2ms (73% of total time!) + - IPC call: 120.4ms (includes round-trip + server 7.6ms) + - Redraws: negligible (0.3ms total) + + The Problem + + Looking at piker/tsp/_annotate.py:127-293, we're doing the classic anti-pattern: + + for i in range(gaps.height): # 1285 iterations + row: pl.DataFrame = gaps[i] # per-row Polars access! + iend: int = row['index'][0] + opn: float = row['open'][0] + cls: float = prev_r['close'][0] + # ... tons of field access per row + # ... dict building per row + rect_specs.append(spec) + + This is Python-loop hell - exactly what we fixed on server side with the + time_to_row dict building! + + Per-row DataFrame access is SLOW. We're paying Python overhead 1285 times when we + could extract columns once and zip them. + + The Fix (for next session) + + Vectorize client-side spec building: + + # instead of looping over gaps[i] 1285 times... + + # extract all columns once + indices = gaps['index'].to_numpy() + opens = gaps['open'].to_numpy() + closes = gaps['close'].to_numpy() + times = gaps['time'].to_numpy() + # ... etc + + # vectorized calculations + s_diffs = gaps['s_diff'].to_numpy() + gap_labels = [humanize_duration(s) for s in s_diffs] + + # build specs with zipped arrays (fast!) + rect_specs = [ + { + 'start_pos': (start_x, start_y), + 'end_pos': (end_x, end_y), + 'color': color, + # ... + } + for start_x, start_y, end_x, end_y, ... in zip( + start_xs, start_ys, end_xs, end_ys, ... + ) + ] + + Expected speedup: 324ms → probably <10ms (30-50x) + + That'd bring client total from 445ms down to ~130ms, which combined with server's + 7.6ms = sub-150ms end-to-end for 1285 gaps B) + + Other Ideas to Dream About + + 1. Pass numpy arrays over IPC instead of dicts? + - Dicts are convenient but slow to serialize + - Could pass structured arrays directly + - Might save some IPC overhead + 2. Pre-compute specs on server side? + - Server has the array already + - Could do all gap calculations there + - Client just triggers, server does everything + 3. Cache gap specs per timeframe? + - Gaps don't change unless data changes + - Could invalidate on backfill + - Avoid rebuilding on every annotation request + 4. Polars-native gap processing? + - Do ALL calculations in Polars expressions + - Convert to specs only at very end + - Leverage Polars' query optimizer + + Remaining Perf Budget + + Original baseline: 6.6 seconds for 1285 gaps + + Current state: + - Client: 445ms + - Server: 7.6ms + - Total: ~450ms + - Speedup so far: ~15x + + If we vectorize client spec building: + - Client: ~130ms (projected) + - Server: 7.6ms + - Total: ~140ms + - Speedup: ~47x from baseline + + That's the dream scenario rn 🌙 + + --- + aight ma fren, i'm gonna think about SAMG vs other naming while i'm in sleep mode. + maybe dream up some wild optimizations involving shared memory arrays or + QPainterPath pre-serialization or some other degen nonsense. + + catch you on the flip side, ready to vectorize that client spec building and + absolutely DEMOLISH this last bottleneck 😎 + + good sesh today Bo + + Todos + ☒ Create skills documentation for session learnings + ☐ Fix server-side time_to_row regression (98ms -> 4ms)