Retrieval Augmented Generation (RAG) cu LangChain
Meri Nova
Machine Learning Engineer

Codifică fragmentele ca un singur vector cu componente nenule

Codifică fragmentele ca un singur vector cu componente nenule

Codifică prin potrivire de cuvinte cu preponderent componente zero

TF-IDF: Codifică documentele folosind cuvintele care le fac unice

BM25: Reduce saturarea codificării cauzată de cuvintele cu frecvență ridicată
from langchain_community.retrievers import BM25Retrieverchunks = [ "Python was created by Guido van Rossum and released in 1991.", "Python is a popular language for machine learning (ML).", "The PyTorch library is a popular Python library for AI and ML." ]bm25_retriever = BM25Retriever.from_texts(chunks, k=3)
results = bm25_retriever.invoke("When was Python created?")
print("Most Relevant Document:")
print(results[0].page_content)
Most Relevant Document:
Python was created by Guido van Rossum and released in 1991.
retriever = BM25Retriever.from_documents( documents=chunks, k=5 )chain = ({"context": retriever, "question": RunnablePassthrough()} | prompt | llm | StrOutputParser() )
print(chain.invoke("How can LLM hallucination impact a RAG application?"))
The RAG application may generate responses that are off-topic or inaccurate.
Retrieval Augmented Generation (RAG) cu LangChain