LangExtract: Python library for extracting structured data from language models

constantinum · 2025-08-03T02:40:55 1754188855

There is also Unstract(open-source) that helps process structured data extraction. Key differences:

1. Unstract has a Pre-processing layer(OCR). Which converts documents into LLM readable formats.(helps improve accuracy, and control costs)

2. Unstract also connects to your existing data sources, making it an out-of-the-box ETL tool.

https://github.com/Zipstack/unstract

ttul · 2025-08-03T16:34:54 1754238894

I’d throw a vote in the column for Unstract. Making the code AGPL is a first class move for a company that is trying to make money from the hosted version of the same software.

fudged71 · 2025-08-03T05:33:31 1754199211

Any idea how it compares with docetl?

oriettaxx · 2025-08-03T03:00:59 1754190059

impressive, really

simonw · 2025-08-03T16:18:43 1754237923

I implemented a similar pattern in my LLM tool and Python library back in February: https://simonwillison.net/2025/Feb/28/llm-schemas/

My version works with Pydantic models or JSON schema in Python code, or with JSON schema or a weird DSL I invented on the command-line:

  curl https://news.ycombinator.com/ | \
    llm --schema-multi 'headline,url,votes int' \
    -m gpt-4.1 --system 'all links'

Result: https://gist.github.com/simonw/f8143836cae0f058f059e1b8fc2d9...

wodenokoto · 2025-08-03T08:47:35 1754210855

It’s not extracting data _from_ the model it is using the model to extract structured data from the input.

ttul · 2025-08-03T16:32:17 1754238737

The use case that immediately comes to mind is analysis of legal documents. Lawyers spend a lot of time going through piles of contracts during due diligence for any kind of investment or acquisition transaction, painstakingly identifying concepts that need to be addressed in various ways. LLMs are decent at doing this kind of work, but error-prone (as are humans, by the way). Having a way to visualize the results could be helpful in speeding up the review process of the LLM’s work.

hm-nah · 2025-08-03T04:37:36 1754195856

Oly Chit! This is a BIG deal! Sub-page citations…in-context RAG…built-in HTML UI…this is like the holy grail of deterministic text extraction. I’m trying this ASAP Rocky.

albert_e · 2025-08-04T04:18:23 1754281103

For complex business documents -- one approach was to use Named Entity Recognition to identify all entities and use that to build a knowledge graph to serve as a complementary repository of knowledge (in addition to the vector embeddings of semantic chunks) to aid RAG workflows.

Does this proposed approach complement this or supercede the need for NER / Knowledge Graph. Just wondering aloud. Appreciate any insights here.

Noumenon72 · 2025-08-03T14:53:51 1754232831

In the example, if `extraction_class` can be any string, how does it know that "relationship" implies it should have attributes "character_1" and "character_2" when your example data didn't?

andrewrn · 2025-08-03T16:16:48 1754237808

You could use this to generate character graphs from big novels. Make an app that allows you to input a page number so the model only extracts characters you've encountered thus far.

ramkumarkb · 2025-08-03T15:09:15 1754233755

Does this work with other open-source LLMs like Qwen3 or other OpenAI compatible LLM Apis?

simonw · 2025-08-03T16:41:06 1754239266

The README says:

> For developers using local LLMs, LangExtract offers built-in support for Ollama and can be extended to other third-party APIs by updating the inference endpoints.

If you look in the code they currently have classes for Gemini and Ollama: https://github.com/google/langextract/blob/main/langextract/...

If you want to do structured data extraction with a wider variety of libraries I'm going to promote my LLM library and tool, which supports dozens of models for this via the plugins mechanism: https://llm.datasette.io/en/stable/schemas.html

brokensegue · 2025-08-03T17:20:12 1754241612

wiring this to wikidata would be great