Building Decision Provenance Into AI-Assisted Data Science Workflows
I built this plugin for a boring reason: decisions in a data science project don't survive.Not nece 2026-10-6 13:40:55 Author: hackernoon.com(查看原文) 阅读量:3 收藏

I built this plugin for a boring reason: decisions in a data science project don't survive.

Not necessarily because nobody writes them down — but because they were never the kind of thing anybody writes down. You tried 80/20 in August; the holdout leaked promotions, you moved on. Someone in a review said "that's a time series", so the split became a rolling window. Upstream stopped shipping a column in v2, so there's handling for it now.

Six months later, none of that exists anywhere. The code shows the outcome and nothing else: a filter with no reason, a ratio with no history, a rejected approach that left no trace of having been rejected. So the next person tries it again.

So does your agent.

That was the problem I set out to fix. Record the decision next to the thing it explains, per task and per node, and put it in front of the agent before it touches that thing again. Stop re-running experiments that already failed. Make every change carry the reason it was made.

Then something else happened.

yzhao's image-e553a

The part I didn't design

I got questioned about a change.

The change wasn't mine. They had asked for it. I had done it, and — because the plugin was running — I had recorded it.

So I typed one command and put the date, their words, and the email on the table.

The meeting moved on.

I built provenance. What I got was receipts.

Fine, let's talk about that

I'm not going to pretend that isn't part of it. git blame tells you who wrote a line. This tells you who asked for it.

But here's the thing I only understood after living with it for a few months: the person who changed their mind doesn't remember either.

They're not hiding it from you. That decision was a two-minute call for them. They've made four hundred more since. The one currently ruining your September barely registered at the time. When you produce the email, the reaction isn't usually defensiveness — it's "oh, right, that."

So the receipts framing is funny, and it's real, but it's not the point. The point is that a group of people who are all individually doing their best are, collectively, forgetting everything. The plugin is a shared memory for that group. It just happens to be admissible.

And practically: a filter with a reason attached survives the next refactor. A filter without one does not. That has nothing to do with anybody's feelings and it's most of the value.

yzhao's image-559d6

What the plugin actually is

It's called provLedger. MIT, no service, no account. Everything it records lives in two SQLite files you can open yourself.

One graph holds your code, your data (dataframes, columns, tables, feeds, metrics), the dependencies between them in both directions, and the reason behind each change recorded next to the thing it changed.

Three things it does.

1 — It stops your agent repeating a failed experiment

> change the train/test split to 80/20 and retrain the churn model

  Planning · 4 steps
  provledger · declaring 3 targets, checking each against its own history
  provledger · computing what reads them — 5 consumers, 3 steps deep

  ⚠ 2 findings before anything is edited

  1  This experiment has been run before
     80/20 was tried on 2026-08-14 and rejected — the holdout leaked
     week-52 promotions, so the lift was the promotion and not the model.

  2  weekly_report depends on this split
     split_dataset → train_churn → weekly_report — the churn figure in
     the Monday deck.

  ↳ nothing is blocked — recorded, and put in front of you

Nothing is blocked. You might have a good reason to try 80/20 again. But you're not going to spend a day rediscovering why it didn't work in August.

2 — It keeps what was actually said

Open the record and you get the whole history of that node, not just this week's:

2026-09-17  Proposed: move the split to 80/20 and retrain.        stated
2026-08-14  Rejected: 80/20 leaked week-52 promotions into the
            holdout. Keep 70/30 until the promotion calendar is
            a column.                                    3 hits · stated

    "We tried eighty-twenty in August and the holdout had week
     fifty-two in it — the lift was the promotion, not the model.
     Leave it at seventy-thirty until the promotion calendar is
     its own column."

    tier          stated — the user's own words, with the span
    source level  verbal — the words themselves
    recorded in   churn-t094, 2026-08-14 11:07
    surfaced      3 times · adopted by 1 plan

2026-06-02  Stratify on churn, not on tenure.                5 hits · stated
2026-05-19  Signature changed: returns (train, test, meta).         derived
2026-04-03  No reason on record for the original 70/30 choice.     unstated

Two things to notice.

None of that is the agent's summary. It's what was said, kept verbatim, with a label saying so. The agent's own interpretation is stored separately and marked as an interpretation, because a tasteful paraphrase is worth nothing when someone asks you what was actually agreed.

That last line is a real answer. unstated means the database checked and there is no reason on record. Not "I couldn't find anything" — an actual, computed absence. Which is how you find out that the 70/30 you've defended twice was never justified by anyone.

3 — It knows where the data actually flows

The graph is rebuilt from source on every run and diffed against the previous one. So "this column is gone", "this output type changed", "this is the same node after a rename" are measured, not reported by a model.

That lets it answer things an AST-only tool can't:

  • this column isn't from the source, it's derived after import
  • this feature reaches back into data the model shouldn't see
  • nineteen things read this intermediate table, here they are

A model fit and evaluated on the same lineage root without a real split gets flagged as an error, not a warning. Because it is one.

Honest drawbacks

It rests on what you say to your agent. If a decision happens in a meeting and nobody ever states it, the ledger will tell you the change is unexplained rather than inventing a reason for it.

It proves a record existed at a point in time and hasn't been altered since. It does not prove the record is true. Those are different claims and I'm not going to blur them.

Nothing that decides an outcome is left to the model. It has two jobs: matching what you said to what changed, and turning a computed table of facts into a sentence. Every number in that sentence is checked against the table afterwards, in code, and anything that cites nothing gets deleted.

Anything model-dependent ships switched off until it clears an accuracy gate. Two features currently haven't. That's in the changelog, not quietly omitted.

Try it

make demo runs the whole thing offline and deterministically in two minutes from a fresh clone. No API key, no data, no account.


文章来源: https://hackernoon.com/building-decision-provenance-into-ai-assisted-data-science-workflows?source=rss
如有侵权请联系:admin#unsafe.sh