Nate Mayer
LX*6~ Teaching AI to Read Payer Edits · Part 6

What Iterating Taught Me About LangChain

Why the payer-edit pipeline moved off LangChain once the task turned out to be four fixed steps rather than an agentic graph.

DTM*405*20261002~

We read Generative AI with LangChain in my book club a while back, and that's where this series started. I built the payer-edit pipeline on LangChain to work through the book's ideas on a real system. It got me from idea to working pipeline fast, and the pipeline did its job, turning companion guides into validation rules that met the pattern standard I was after.

I've spent the time since iterating, and mostly learning how to set the whole thing up better.

The biggest thing I learned was the shape of the problem. Turning a payer edit into a validation rule isn't an open-ended, agentic task, because every edit goes through the same four steps in the same order:

  1. Classify and decompose the edit to work out what it's actually validating.
  2. Map it to X12 segments and elements.
  3. Generate the rule logic.
  4. Generate the human-facing rejection message.

If a generated rule fails validation, the pipeline turns the errors into guidance and runs the four steps once more, and if the rule still fails, it stops there. The order never changes, there are no tool calls, and no conversation carries over from one call to the next, because the task doesn't need any of that.

Once I understood that, the rest followed, starting with the flow around the four steps. That flow ran on LangGraph, which is made for workflows that loop, wait for a person, or pick up again after a restart. Mine picked a parser for the document's format, turned the parsed document into an edit, and retried a failed rule once, and plain Python does all three. LangGraph's checkpoints let a run resume partway through a document, but the batch already records each edit as it finishes, so I didn't need a second way to pick up where it left off.

I also cared about prompt caching and about checking each step's output, and LangChain made both of them indirect. Caching only pays off when every call starts with exactly the same text. I could mark which blocks to cache through LangChain, but LangChain assembled the final request, so whether those blocks stayed identical from call to call depended on library code I didn't own.

The JSON parser my steps used returned a plain dictionary, so only step 3's output got checked, and only for required fields, by code I wrote. LangChain has parsers that return Pydantic models, but with the raw SDK, checking a step takes one line, because Pydantic either accepts the output or says exactly what's wrong with it.

So I rebuilt it on the raw Anthropic and OpenAI SDKs, with plain Python orchestrating the four steps, Pydantic validating each one, and no LangChain in between. The control flow is now the four steps, written out in order, so I can read the whole thing top to bottom. The trade is that the code that talks to each provider is now a small client I own, where LangChain used to supply those adapters.