Metodi avanzati di suddivisione

Retrieval Augmented Generation (RAG) con LangChain

Meri Nova

Machine Learning Engineer

Limiti delle strategie di suddivisione attuali

 

  1. 🤦 Le suddivisioni sono ingenue (non considerano il contesto)

    • Ignorano il contesto del testo circostante
  2. 🖇 La divisione usa caratteri invece di token

    • I token sono elaborati dai modelli
    • Rischio di superare la context window

 

SemanticChunker

 

TokenTextSplitter

Retrieval Augmented Generation (RAG) con LangChain

Divisione sui token

Uno splitter per caratteri che divide il testo in blocchi in base al numero di caratteri.

Retrieval Augmented Generation (RAG) con LangChain

Divisione sui token

Uno splitter per token che divide il testo in blocchi in base al numero di token.

Retrieval Augmented Generation (RAG) con LangChain

Divisione sui token

I token sono evidenziati per mostrare l'allineamento con i valori di chunk_size e chunk_overlap.

Retrieval Augmented Generation (RAG) con LangChain

Divisione sui token

import tiktoken
from langchain_text_splitters import TokenTextSplitter
example_string = "Mary had a little lamb, it's fleece was white as snow."

encoding = tiktoken.encoding_for_model('gpt-4o-mini')
splitter = TokenTextSplitter(encoding_name=encoding.name,
                             chunk_size=10,
                             chunk_overlap=2)

chunks = splitter.split_text(example_string) for i, chunk in enumerate(chunks): print(f"Chunk {i+1}:\n{chunk}\n")
Retrieval Augmented Generation (RAG) con LangChain

Divisione sui token

Chunk 1:
Mary had a little lamb, it's fleece

Chunk 2:
 fleece was white as snow.
Retrieval Augmented Generation (RAG) con LangChain

Divisione sui token

for i, chunk in enumerate(chunks):
    print(f"Chunk {i+1}:\nNo. tokens: {len(encoding.encode(chunk))}\n{chunk}\n")
Chunk 1:
No. tokens: 10
Mary had a little lamb, it's fleece was

Chunk 2:
No. tokens: 6
 fleece was white as snow.
Retrieval Augmented Generation (RAG) con LangChain

Suddivisione semantica

Un paragrafo con una frase sulle applicazioni RAG e una sui cani.

Retrieval Augmented Generation (RAG) con LangChain

Suddivisione semantica

Il paragrafo è stato diviso usando caratteri o token, perdendo il contesto.

Retrieval Augmented Generation (RAG) con LangChain

Suddivisione semantica

Uno splitter semantico ha diviso il paragrafo nel punto in cui l'argomento passa da RAG ai cani.

Retrieval Augmented Generation (RAG) con LangChain

Suddivisione semantica

from langchain_openai import OpenAIEmbeddings
from langchain_experimental.text_splitter import SemanticChunker

embeddings = OpenAIEmbeddings(api_key="...", model='text-embedding-3-small')
semantic_splitter = SemanticChunker( embeddings=embeddings,
breakpoint_threshold_type="gradient", breakpoint_threshold_amount=0.8
)
1 https://api.python.langchain.com/en/latest/text_splitter/langchain_experimental.text_splitter. SemanticChunker.html
Retrieval Augmented Generation (RAG) con LangChain

Suddivisione semantica

chunks = semantic_splitter.split_documents(data)
print(chunks[0])
page_content='Retrieval-Augmented Generation for\nKnowledge-Intensive NLP Tasks\ Patrick Lewis,
Ethan Perez,\nAleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich
Küttler,\nMike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, Douwe Kiela\nFacebook AI
Research; University College London;New York University;\[email protected]\nAbstract\nLarge
pre-trained language models have been shown to store factual knowledge\nin their parameters,
and achieve state-of-the-art results when fine-tuned on down-\nstream NLP tasks. However, their
ability to access and precisely manipulate knowl-\nedge is still limited, and hence on
knowledge-intensive tasks, their performance\nlags behind task-specific architectures.'
metadata={'source': 'rag_paper.pdf', 'page': 0}
Retrieval Augmented Generation (RAG) con LangChain

Esercitiamoci!

Retrieval Augmented Generation (RAG) con LangChain

Preparing Video For Download...