Méthodes avancées de découpage

Retrieval Augmented Generation (RAG) avec LangChain

Meri Nova

Machine Learning Engineer

Limites de nos stratégies actuelles de découpage

 

  1. 🤦 Découpage naïf (sans tenir compte du contexte)

    • Ignore le contexte du texte autour
  2. 🖇 Découpage par caractères plutôt que par jetons

    • Les jetons sont traités par les modèles
    • Risque de dépasser la fenêtre de contexte

 

SemanticChunker

 

TokenTextSplitter

Retrieval Augmented Generation (RAG) avec LangChain

Découper par jetons

Un séparateur par caractères découpe le texte en blocs selon le nombre de caractères.

Retrieval Augmented Generation (RAG) avec LangChain

Découper par jetons

Un séparateur par jetons découpe le texte en blocs selon le nombre de jetons.

Retrieval Augmented Generation (RAG) avec LangChain

Découper par jetons

Les jetons sont mis en évidence pour montrer leur alignement avec `chunk_size` et `chunk_overlap`.

Retrieval Augmented Generation (RAG) avec LangChain

Découper par jetons

import tiktoken
from langchain_text_splitters import TokenTextSplitter
example_string = "Mary had a little lamb, it's fleece was white as snow."

encoding = tiktoken.encoding_for_model('gpt-4o-mini')
splitter = TokenTextSplitter(encoding_name=encoding.name,
                             chunk_size=10,
                             chunk_overlap=2)

chunks = splitter.split_text(example_string) for i, chunk in enumerate(chunks): print(f"Chunk {i+1}:\n{chunk}\n")
Retrieval Augmented Generation (RAG) avec LangChain

Découper par jetons

Chunk 1:
Mary had a little lamb, it's fleece

Chunk 2:
 fleece was white as snow.
Retrieval Augmented Generation (RAG) avec LangChain

Découper par jetons

for i, chunk in enumerate(chunks):
    print(f"Chunk {i+1}:\nNo. tokens: {len(encoding.encode(chunk))}\n{chunk}\n")
Chunk 1:
No. tokens: 10
Mary had a little lamb, it's fleece was

Chunk 2:
No. tokens: 6
 fleece was white as snow.
Retrieval Augmented Generation (RAG) avec LangChain

Découpage sémantique

Un paragraphe contenant une phrase sur les applications RAG et une phrase sur les chiens.

Retrieval Augmented Generation (RAG) avec LangChain

Découpage sémantique

Le paragraphe a été découpé par caractères ou jetons, ce qui a fait perdre le contexte.

Retrieval Augmented Generation (RAG) avec LangChain

Découpage sémantique

Un séparateur sémantique a scindé le paragraphe là où le sujet passe de RAG aux chiens.

Retrieval Augmented Generation (RAG) avec LangChain

Découpage sémantique

from langchain_openai import OpenAIEmbeddings
from langchain_experimental.text_splitter import SemanticChunker

embeddings = OpenAIEmbeddings(api_key="...", model='text-embedding-3-small')
semantic_splitter = SemanticChunker( embeddings=embeddings,
breakpoint_threshold_type="gradient", breakpoint_threshold_amount=0.8
)
1 https://api.python.langchain.com/en/latest/text_splitter/langchain_experimental.text_splitter. SemanticChunker.html
Retrieval Augmented Generation (RAG) avec LangChain

Découpage sémantique

chunks = semantic_splitter.split_documents(data)
print(chunks[0])
page_content='Retrieval-Augmented Generation for\nKnowledge-Intensive NLP Tasks\ Patrick Lewis,
Ethan Perez,\nAleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich
Küttler,\nMike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, Douwe Kiela\nFacebook AI
Research; University College London;New York University;\[email protected]\nAbstract\nLarge
pre-trained language models have been shown to store factual knowledge\nin their parameters,
and achieve state-of-the-art results when fine-tuned on down-\nstream NLP tasks. However, their
ability to access and precisely manipulate knowl-\nedge is still limited, and hence on
knowledge-intensive tasks, their performance\nlags behind task-specific architectures.'
metadata={'source': 'rag_paper.pdf', 'page': 0}
Retrieval Augmented Generation (RAG) avec LangChain

Passons à la pratique !

Retrieval Augmented Generation (RAG) avec LangChain

Preparing Video For Download...