← Writing
AI & Generative AI

LLM Note 4

Data Mastery Series — Episode 49: Understanding RAG Taxonomy

22 Feb 202545 min readLangChainLangGraphRAGAI AgentDashboard
LLM Note · Part 4 of 5

LLM Note 4

Data Mastery Series — Episode 49: Understanding RAG Taxonomy

Connect with me and follow our journey: Linkedin, Facebook

1) A Taxonomy of Retrieval Augmented Generation

(SOURCE)เนื้อหาใน Blog จะแบ่งหัวข้อออกเป็น 8 ส่วนหลัก เพื่ออธิบายองค์ประกอบต่างๆ ของ RAG อย่างละเอียด:

Retrieval Augmented Generation enhances the reliability and the trustworthiness in LLM responses (Source: A Taxonomy of Retrieval Augmented Generation)

  1. RAG Basics (พื้นฐานของ RAG): อธิบายข้อจำกัดของ LLM และแนะนำแนวคิด RAG เบื้องต้น
  2. Core Components (องค์ประกอบหลัก): อธิบายกระบวนการสร้าง index และ generation
  3. Evaluation (การประเมินผล): อธิบาย metrics และ frameworks ที่ใช้ในการประเมินประสิทธิภาพของ RAG
  4. Pipeline Design (การออกแบบ pipeline): อธิบายรูปแบบต่างๆ ของ RAG pipelines (เช่น แบบง่าย, แบบซับซ้อน, แบบ modular)
  5. Operations Stack (Operations Stack): อธิบายเครื่องมือและเทคนิคที่ใช้ในการจัดการและ monitor RAG systems
  6. Emerging Patterns (รูปแบบใหม่ๆ): อธิบายแนวทาง RAG ที่กำลังพัฒนา เช่น multimodal RAG, agentic RAG และ KG-powered RAG
  7. Technology Providers (ผู้ให้บริการเทคโนโลยี): รายชื่อโซลูชันและผู้ให้บริการที่เกี่ยวข้องกับ RAG
  8. Applied RAG (RAG ในการใช้งานจริง): รูปแบบการใช้งาน RAG, ขอบเขตการใช้งาน, และความท้าทายในปัจจุบัน

RAG Basics (พื้นฐานของ RAG)

1️⃣ LLM Limitations (ข้อจำกัดของ LLM)

  • Knowledge Cut-off Date (ข้อมูลอัพเดตถึงเมื่อไร):
    LLMs ถูก train บนข้อมูลจำนวนมหาศาล แต่ข้อมูลนั้นไม่ได้เป็นปัจจุบัน
  • Training Data Limitation (ข้อจำกัดของข้อมูลที่ใช้ train):
    LLMs ถูก train บนข้อมูลสาธารณะเท่านั้น
  • Hallucinations (การสร้างข้อมูลเท็จ):
    LLMs อาจให้ข้อมูลที่ไม่ถูกต้อง แต่ฟังดูน่าเชื่อถือ
  • Context Window (ความยาวของคำตอบ):
    คำตอบของ LLMs ถูกจำกัดตามจำนวน tokens (หน่วยของคำ) ถ้า prompt ยาวเกิน context window ส่วนเกินจะถูกตัดทิ้ง

2️⃣ RAG Concepts (แนวคิด RAG)

  • Parametric Memory (หน่วยความจำเชิงพารามิเตอร์):
    ความรู้ที่ฝังอยู่ในตัว LLM เอง, ได้มาจากการ train (การสอน) เช่น 1.5 trillion parameters ของ Gemini 1.5 Pro ซึ่งเปลี่ยนแปลงไม่ได้ง่ายๆ
  • Non-parametric Memory (หน่วยความจำที่ไม่ใช่เชิงพารามิเตอร์):
    ข้อมูลที่ LLM ไม่มีอยู่ใน parameters แต่สามารถเข้าถึงได้จากภายนอก เช่น search engine หรือ database RAG ซึ่งสามารถเปลี่ยนแปลงและอัพเดทได้
  • Knowledge Base (ฐานความรู้):
    แหล่งข้อมูลภายนอกที่สร้างขึ้นสำหรับ RAG application ข้อมูลถูกประมวลผลและจัดเก็บใน persistent memory (เช่นเก็บใน Database สำหรับ RAG)
  • User Query (คำถามของผู้ใช้):
    Prompt ที่ผู้ใช้ส่งไปยัง LLM เพื่อขอคำตอบ
  • Retrieval (การดึงข้อมูล):
    กระบวนการค้นหาและดึงข้อมูลที่เกี่ยวข้องกับ user query จาก knowledge base
  • Augmentation (การเสริมข้อมูล):
    กระบวนการเพิ่มข้อมูลที่ถูกดึงมาลงใน user query
  • Generation (การสร้างข้อความ):
    กระบวนการสร้างผลลัพธ์โดย LLM โดยใช้ augmented prompt
  • Source Citation (การอ้างอิงแหล่งที่มา):
    ความสามารถของ RAG system ในการชี้ไปยังข้อมูลจาก knowledge base ที่ใช้ในการสร้างคำตอบ
  • Unlimited Memory (หน่วยความจำไม่จำกัด):
    แนวคิดที่ว่าสามารถเพิ่ม documents จำนวนเท่าใดก็ได้ลงใน RAG knowledge base

How does RAG help? (Source: A Taxonomy of Retrieval Augmented Generation)

Core Components (องค์ประกอบหลัก)

1️⃣ Indexing (การสร้างดัชนี)

  • Indexing Pipeline (ขั้นตอนการสร้างดัชนี):
    ชุดของกระบวนการที่ใช้ในการสร้าง knowledge base สำหรับ RAG applications เป็น pipeline ที่ไม่ใช่ real-time และจะ update knowledge base เป็นช่วงๆ
  • Source Systems (ระบบต้นทาง):
    ที่เก็บข้อมูลดั้งเดิมที่ต้องการใช้สำหรับ RAG application เช่น data lakes, file systems, CMSs, SQL & NoSQL databases
  • Data Loading (การโหลดข้อมูล):
    ขั้นตอนแรกของ indexing pipeline ที่เชื่อมต่อกับ source systems เพื่อดึงและ parse files เพื่อนำข้อมูลไปใช้ใน RAG knowledge base
  • Metadata (ข้อมูลอธิบายข้อมูล):
    ข้อมูลที่อธิบายข้อมูลอื่นๆ ให้ข้อมูล เช่น รายละเอียดของข้อมูล, เวลาที่สร้าง, ผู้สร้าง Metadata ช่วยให้ค้นหาข้อมูลได้ง่ายขึ้นใน RAG
  • Data Masking (การปกปิดข้อมูล):
    การปิดบังข้อมูลที่ละเอียดอ่อน เช่น PII (Personally Identifiable Information) หรือข้อมูลที่เป็นความลับ
  • Chunking (การแบ่งข้อมูลเป็นชิ้น):
    กระบวนการแบ่งข้อความยาวๆ ออกเป็นชิ้นเล็กๆ เพื่อให้จัดการได้ง่ายขึ้น Chunking ช่วยให้ค้นหาได้ง่ายขึ้น และแก้ปัญหา context window limits ของ LLMs
  • Lost in the middle problem (ปัญหาการหลงลืมตรงกลาง):
    LLMs อาจอ่านข้อมูลได้ไม่แม่นยำถ้าข้อมูลที่เกี่ยวข้องอยู่ตรงกลางของ prompt การส่งเฉพาะข้อมูลที่เกี่ยวข้องไปยัง LLM สามารถแก้ปัญหานี้ได้
  • Fixed Size Chunking (การแบ่งข้อมูลเป็นชิ้นขนาดคงที่):
    กำหนดขนาดของ chunk และ overlap (ส่วนที่ซ้อนทับกัน) ล่วงหน้า
  • Structure-Based Chunking (การแบ่งข้อมูลตามโครงสร้าง):
    แบ่งข้อมูลตามโครงสร้าง เช่น HTML, Markdown, JSON แทนที่จะใช้ขนาดคงที่
  • Context-Enriched Chunking (การแบ่งข้อมูลเสริมบริบท):
    เพิ่ม summary ของ document ที่ใหญ่กว่าลงในแต่ละ chunk เพื่อเสริม context
  • Agentic Chunking (การแบ่งข้อมูลแบบ Agent):
    สร้าง chunks โดยอิงตามเป้าหมายหรือ task เช่น วิเคราะห์ customer reviews โดยแบ่ง reviews ตาม topic หรือ sentiment
  • Semantic Chunking (การแบ่งข้อมูลเชิงความหมาย):
    แบ่งข้อมูลโดยพิจารณา semantic similarity (ความคล้ายคลึงกันในความหมาย) ระหว่าง sentences
  • Small to big chunking (การแบ่งข้อมูลจากเล็กไปใหญ่):
    แบ่งข้อความเป็นหน่วยเล็กๆ ก่อน (เช่น sentences, paragraphs) แล้วรวมหน่วยเล็กๆ เข้าด้วยกันจนได้ขนาด chunk ที่ต้องการ

Small to big chunking (Source: A Taxonomy of Retrieval Augmented Generation)

  • Chunk Size (ขนาดของชิ้นข้อมูล):
    ขนาดของ chunk มีผลต่อคุณภาพของ RAG system chunks ขนาดใหญ่ให้ context ที่ดีกว่า แต่ก็มี noise มากกว่า chunks ขนาดเล็กมีข้อมูลที่แม่นยำกว่า แต่อาจพลาดข้อมูลสำคัญ
  • Metadata Filtering (การกรองด้วย metadata):
    เพิ่ม metadata เช่น timestamp, author, category เพื่อกรอง chunks ก่อนทำ similarity search
  • Metadata Enhancement (การเพิ่มประสิทธิภาพ metadata):
    ใช้ metadata ที่ inferred (อนุมาน) ได้ เช่น chunk summary, sentiment, category เพื่อเพิ่มประสิทธิภาพการ retrieval
  • Parent Child Indexing (การสร้างดัชนีแบบ Parent-Child):
    จัดระเบียบ documents ในลักษณะ hierarchical parent document มี overarching themes (หัวข้อหลัก) child documents เจาะลึกรายละเอียด ระบบจะ locate child documents ที่เกี่ยวข้องก่อน แล้วอ้างอิง parent documents เพื่อ context เพิ่มเติม
  • Embeddings (การแปลงข้อมูลเป็นเวกเตอร์):
    การแปลงข้อมูลที่ไม่ใช่ตัวเลข เช่น text หรือ image ให้เป็นรูปแบบตัวเลข (vectors) Embeddings ใช้สำหรับ RAG เพราะช่วยในการสร้าง semantic relationship (ความสัมพันธ์เชิงความหมาย) ระหว่างคำ, phrases และ documents
  • Cosine Similarity (ความคล้ายคลึงโคไซน์):
    คำนวณความคล้ายคลึงกันระหว่าง vectors โดยวัด cosine ของมุมระหว่าง vectors Terms ที่เกี่ยวข้องจะมี cosine similarity ใกล้ 1 terms ที่ไม่เกี่ยวข้องจะมีค่าใกล้ 0
  • Word2Vec:
    โมเดล shallow neural network ที่ใช้เรียนรู้ word embeddings พัฒนาโดย Google
  • GloVe:
    Unsupervised learning technique ที่พัฒนาโดย Stanford University
  • FastText:
    ส่วนขยายของ Word2Vec พัฒนาโดย Facebook AI Research มีประโยชน์สำหรับการจัดการ misspellings และ rare words
  • ELMo:
    Embeddings from Language Models พัฒนาโดย Allen Institute for AI
  • BERT:
    Bidirectional Encoder Representations from Transformers พัฒนาโดย Google เป็นโมเดลที่ใช้ architecture แบบ Transformers ให้ contextualized word embeddings
  • Pre-trained Embeddings Models:
    Embeddings models ที่ถูก train บนข้อมูลจำนวนมาก สามารถ generalize ได้ดีกับ tasks และ domains ต่างๆ
  • Vector Databases (ฐานข้อมูลเวกเตอร์):
    สร้างขึ้นเพื่อจัดการ high dimensional vectors เช่น embeddings ฐานข้อมูลเหล่านี้ specialize ในการ indexing และจัดเก็บ vector embeddings เพื่อให้ค้นหา semantic search และ retrieval ได้อย่างรวดเร็ว
  • Vector Indices (ดัชนีเวกเตอร์):
    Libraries ที่เน้น core features ของ indexing และ search ไม่ support data management, query processing, interfaces

2️⃣ Generation (การสร้างข้อความ)

  • Generation Pipeline (ขั้นตอนการสร้างข้อความ):
    ชุดของกระบวนการที่ใช้ในการค้นหาและดึงข้อมูลจาก knowledge base เพื่อสร้าง responses ต่อ user queries
  • Information Retrieval (IR) (การดึงข้อมูล):
    ศาสตร์แห่งการค้นหา ไม่ว่าจะเป็นการค้นหาข้อมูลใน document หรือค้นหา documents เอง
  • Retriever (ตัวดึงข้อมูล):
    องค์ประกอบของ generation pipeline ที่ใช้อัลกอริทึมในการค้นหาและดึงข้อมูลที่เกี่ยวข้องจาก knowledge base
  • Boolean retrieval (การดึงข้อมูลแบบ Boolean):
    การค้นหาแบบ keyword ที่ใช้ Boolean logic เพื่อ match documents กับ queries โดยอิงตาม absence หรือ presence ของคำ
  • TF-IDF (Term Frequency-Inverse Document Frequency):
    Statistical measure ที่ใช้ประเมินความสำคัญของคำใน document เมื่อเทียบกับ collection of documents (corpus)

TF-IDF calculation (Source: A Taxonomy of Retrieval Augmented Generation)

  • BM25 (Best Match 25):
    Advanced probabilistic model ที่ใช้ rank documents โดยอิงตาม query terms ที่ปรากฏในแต่ละ document ปรับแก้ความยาวของ documents เพื่อไม่ให้ documents ที่ยาวกว่าได้คะแนนสูงกว่าอย่างไม่เป็นธรรม
  • Static Word Embeddings (Embeddings คำแบบคงที่):
    Embeddings เช่น Word2Vec และ GloVe ที่ represent คำเป็น dense vectors ใน continuous vector space capturing semantic relationships
  • Contextual Embeddings (Embeddings คำตามบริบท):
    Embeddings ที่ generate โดย models เช่น BERT หรือ OpenAI’s text embeddings ให้ high-dimensional, context-aware representations

Static vs Contextual Embeddings (Source: A Taxonomy of Retrieval Augmented Generation)

  • Learned Sparse Retrieval (การดึงข้อมูลแบบ Sparse ที่เรียนรู้ได้): Generate sparse representations โดยใช้ neural networks
  • Dense Retrieval (การดึงข้อมูลแบบ Dense):
    Encode queries และ documents เป็น dense vectors
  • Hybrid Retrieval (การดึงข้อมูลแบบ Hybrid):
    รวม sparse และ dense methods
  • Cross-Encoder Retrieval (การดึงข้อมูลแบบ Cross-Encoder):
    เปรียบเทียบ query-document pairs โดยใช้ transformer models
  • Graph-based Retrieval (การดึงข้อมูลแบบ Graph):
    ใช้ graph structures เพื่อ model relationships ระหว่าง documents
  • Quantum-inspired Retrieval (การดึงข้อมูลแบบ Quantum):
    ใช้วิธีการของ quantum computing ในการดึงข้อมูล
  • Neural IR models (โมเดล IR แบบ Neural):
    ใช้วิธีการของ neural network ในการดึงข้อมูล
  • Augmentation (การเสริมข้อมูล):
    กระบวนการรวม user query และ documents ที่ดึงมาจาก knowledge base
  • Prompt Engineering:
    เทคนิคในการให้ instructions แก่ LLM เพื่อให้ได้ผลลัพธ์ที่ต้องการ สร้าง prompts เพื่อให้ LLM สร้าง responses ที่ถูกต้องและเกี่ยวข้อง
  • Contextual Prompting (การใช้พรอมต์ตามบริบท):
    เพิ่ม instruction เช่น “Answer only based on the context provided.” เพื่อให้ LLM focus เฉพาะข้อมูลที่ให้มา
  • Controlled Generation Prompting (การใช้พรอมต์ควบคุมการสร้างข้อความ):
    เพิ่ม instruction เช่น “If the question cannot be answered based on the provided context, say I don’t know.”
  • Few Shot Prompting (การใช้พรอมต์แบบตัวอย่างน้อย):
    ให้ examples ใน prompt เพื่อ guide การ generation ในแบบที่ต้องการ
  • Chain of Thought Prompting (การใช้พรอมต์แบบลูกโซ่ความคิด):
    เพิ่ม intermediate “reasoning” steps เพื่อปรับปรุง performance ของ LLMs ใน tasks ที่ต้องการ complex reasoning
  • Self Consistency (ความสอดคล้องในตัวเอง):
    Sample multiple reasoning paths และใช้ generations ของแต่ละ path เพื่อให้ได้คำตอบที่สอดคล้องกันมากที่สุด
  • Generated Knowledge Prompting (การใช้พรอมต์แบบสร้างความรู้):
    สร้าง knowledge chains โดยใช้ latent knowledge ของ models เพื่อเสริมสร้าง reasoning
  • Tree of Thoughts Prompting (ToT) (การใช้พรอมต์แบบต้นไม้ความคิด): เทคนิคการ prompt ที่ LLM สร้างความคิดที่เป็นไปได้หลายทาง (เหมือนกิ่งก้านของต้นไม้) แล้วประเมินและเลือกความคิดที่ดีที่สุดเพื่อแก้ปัญหา มีจุดเด่นคือ เหมาะกับปัญหาที่ซับซ้อนที่ต้องการการคิดหลายขั้นตอนและการตัดสินใจ
    Note: Chain-of-Thought (CoT) สร้างแค่เส้นทางความคิดเดียว แต่ ToT สร้างหลายเส้นทาง
  • Automatic Reasoning and Tool-use (ART) (การให้เหตุผลอัตโนมัติและการใช้เครื่องมือ):
    Framework ที่ LLM สามารถใช้เครื่องมือภายนอก (เช่น search engine, API) เพื่อช่วยในการให้เหตุผลและแก้ปัญหาที่ซับซ้อน มีจุดเด่นคือ ช่วยให้ LLM สามารถเข้าถึงข้อมูลและดำเนินการที่อยู่นอกเหนือความสามารถของตัวเอง
    Note: LLM ทั่วไปจำกัดอยู่แค่ข้อมูลที่ตัวเองมี แต่ ART สามารถใช้เครื่องมือช่วยในการหาข้อมูลและตัดสินใจ
  • Automatic Prompt Engineer (APE):
    Framework ที่ LLM ใช้เพื่อสร้างและเลือก prompt ที่ดีที่สุดสำหรับ task นั้นๆ โดยอัตโนมัติ มีจุดเด่นคือ ช่วยลดภาระในการออกแบบ prompt ด้วยมือ และช่วยให้ได้ prompt ที่มีประสิทธิภาพมากยิ่งขึ้น
  • Active Prompt:
    เทคนิคที่ปรับปรุง Chain-of-Thought โดยการปรับ prompt ให้เข้ากับ task นั้นๆ แบบไดนามิก (ปรับเปลี่ยนตามสถานการณ์)
  • ReAct Prompting:
    เทคนิคที่รวม LLM สำหรับการให้เหตุผลและการดำเนินการ (Action) ไปพร้อมๆ กัน มีจุดเด่นคือ LLM สามารถใช้เครื่องมือภายนอกเพื่อหาข้อมูลและดำเนินการได้ และยังสามารถให้เหตุผลเกี่ยวกับข้อมูลที่ได้มา
  • Recursive Prompting (การใช้พรอมต์แบบเรียกซ้ำ):
    เทคนิคการแก้ปัญหาที่ซับซ้อนโดยการแบ่งออกเป็นปัญหาย่อยๆ แล้วใช้ prompt ในการแก้ปัญหาย่อยๆ เหล่านั้นทีละขั้นตอน ผลลัพธ์จากขั้นตอนก่อนหน้าจะถูกนำไปใช้ในขั้นตอนถัดไป มีจุดเด่นคือ เหมาะกับปัญหาที่ต้องใช้การคิดเชิงประกอบ (compositional generalization) เช่น โจทย์คณิตศาสตร์ หรือการตอบคำถามที่ต้องวิเคราะห์หลายขั้นตอน
  • Foundation Models (แบบจำลองพื้นฐาน):
    LLM ขนาดใหญ่ที่ถูก train บนข้อมูลจำนวนมหาศาล มักถูกนำไป fine-tune เพื่อใช้งานเฉพาะทาง ตัวอย่าง: GPT-4o, Gemini 1.5 Pro etc.
  • Supervised Fine-Tuning (SFT) (การปรับแต่งแบบ Supervised):
    การนำ Foundation Model ที่ train ไว้แล้ว มา train เพิ่มเติมบน labeled dataset (ชุดข้อมูลที่มีการระบุคำตอบที่ถูกต้อง) เพื่อให้ model เก่งในงานเฉพาะทาง เช่น การตอบคำถาม หรือการสร้าง chatbot

Supervised Fine-tuning (SFT) of an LLM (Source: A Taxonomy of Retrieval Augmented Generation)

  • Small Language Models (SLMs) (แบบจำลองภาษาขนาดเล็ก):
    LLM ที่มีขนาดเล็กกว่า (parameters น้อยกว่า) Foundation Models

Evaluation (การประเมินผล)

1️⃣ Metrics (ตัวชี้วัด)

  • Evaluation Metrics (ตัวชี้วัดการประเมินผล):
    Quantitative measures (มาตรวัดเชิงปริมาณ) ที่ใช้ประเมิน performance ของ retrieval & generation และโดยรวมของ RAG system

Precision & Recall (Source: A Taxonomy of Retrieval Augmented Generation)

  • Accuracy (ความแม่นยำ):
    สัดส่วนของการทำนายที่ถูกต้อง (ทั้งถูกและไม่ถูก) จากทั้งหมด
    Note: ในฐานข้อมูลขนาดใหญ่ ส่วนใหญ่ documents ไม่เกี่ยวข้องกับ query ทำให้ accuracy อาจสูงเกินจริงและ misleading
  • Precision (ความเที่ยงตรง):
    สัดส่วนของ documents ที่ retrieved มา ที่เกี่ยวข้องกับ query จริงๆ หรือเป็นคำตอบของคำถามที่ว่า “จาก documents ที่ retrieved มาทั้งหมด มีกี่ documents ที่เกี่ยวข้องจริงๆ?” ซึ่งเป็น Metrics ที่เน้นความถูกต้องของผลลัพธ์ที่ได้
  • Precision@k (ความเที่ยงตรงที่ k):
    Precision ที่วัดเฉพาะใน documents ที่ retrieved มา k อันดับแรก
  • Recall (ความครบถ้วน):
    สัดส่วนของ documents ที่เกี่ยวข้องทั้งหมดในฐานข้อมูล ที่ถูก retrieved มา หรือเป็นคำตอบของคำถามที่ว่า “จาก documents ที่เกี่ยวข้องทั้งหมด มีกี่ documents ที่ถูก retrieved มาจริงๆ?” ซึ่งเน้นความครบถ้วนของ documents ที่ retrieved มา
  • F1-score:
    ค่าเฉลี่ย harmonic ของ Precision และ Recall ซึ่งเน้นความ balance ทั้ง Precision และ Recall
  • Mean Reciprocal Rank (MRR) (ค่าเฉลี่ยของส่วนกลับของอันดับ):
    ค่าเฉลี่ยของส่วนกลับ (reciprocal) ของอันดับ (rank) ของผลลัพธ์ที่เกี่ยวข้อง (relevant) อันแรก ในแต่ละ query มีจุดเด่นคือ ให้ความสำคัญกับอันดับของผลลัพธ์ที่เกี่ยวข้องอันแรก หากผลลัพธ์ที่เกี่ยวข้องอันแรกอยู่ในอันดับต้นๆ จะได้คะแนนสูง
  • Mean Average Precision (MAP) (ค่าเฉลี่ยของความเที่ยงตรงเฉลี่ย):
    ค่าเฉลี่ยของ Average Precision (AP) ในแต่ละ query โดย AP คือค่าที่รวม Precision และ Recall ที่ cut-off ต่างๆ (k อันดับแรก) มีจุดเด่นคือ พิจารณาทั้งความถูกต้อง (Precision) และความครบถ้วน (Recall) ของผลลัพธ์ที่ได้จากหลายๆ อันดับ
  • Normalised Discounted Cumulative Gain (nDCG):
    วัดคุณภาพการจัดอันดับ โดยให้คะแนนสูงแก่ผลลัพธ์ที่เกี่ยวข้องที่อยู่ในอันดับต้นๆ และให้คะแนนลดหลั่นลงไปตามอันดับที่ต่ำลง (discounted) แล้ว normalize ค่า มีจุดเด่นคือ สามารถจัดการกับสถานการณ์ที่ผลลัพธ์มีความเกี่ยวข้องในระดับต่างๆ กันได้

Calculating nDCG (Source: A Taxonomy of Retrieval Augmented Generation)

  • Context relevance (ความเกี่ยวข้องของบริบท):
    ประเมินว่า documents ที่ retrieved มาเกี่ยวข้องกับ query เดิมมากน้อยแค่ไหน โดยดูที่ topical alignment (ความสอดคล้องของหัวข้อ), information usefulness (ประโยชน์ของข้อมูล) และ redundancy (ความซ้ำซ้อน) โดยมีหลักการว่า retrieved context ควรมีข้อมูลที่เกี่ยวข้องกับ query เท่านั้น
  • Answer Faithfulness (ความน่าเชื่อถือของคำตอบ):
    วัดว่าคำตอบที่สร้างขึ้นมีความถูกต้องตามข้อเท็จจริง (factually grounded) ใน retrieved context มากน้อยแค่ไหน โดยมีหลักการว่า ข้อเท็จจริงในคำตอบต้องไม่ขัดแย้งกับ context และสามารถ traced back to the source (อ้างอิงแหล่งที่มาได้)
  • Hallucination Rate (อัตราการสร้างข้อมูลเท็จ):
    สัดส่วนของ generated claims (ข้อความที่สร้างขึ้น) ในคำตอบที่ไม่อยู่ใน retrieved context โดยวัดว่า LLM สร้างข้อมูลที่ไม่เกี่ยวข้องกับข้อมูลที่ retrieved มามากน้อยแค่ไหน
  • Coverage (ความครอบคลุม):
    วัดว่าข้อมูลที่เกี่ยวข้องจาก retrieved passages ถูกรวมอยู่ในคำตอบมากน้อยแค่ไหน โดยมีหลักการว่า คำตอบควรครอบคลุมข้อมูลที่สำคัญจาก retrieved passages
  • Answer Relevance (ความเกี่ยวข้องของคำตอบ):
    วัดว่าคำตอบที่สร้างขึ้นมีความเกี่ยวข้องกับ query มากน้อยแค่ไหน โดยดูที่ system’s ability to comprehend the query (ความสามารถในการเข้าใจ query), response being pertinent to the query (คำตอบที่ตรงประเด็น) และ completeness of the response (ความสมบูรณ์ของคำตอบ)
  • Ground truth (ข้อมูลที่แท้จริง):
    ข้อมูลที่รู้ว่าเป็นจริง ใช้เป็น benchmark (เกณฑ์มาตรฐาน) ในการประเมิน
  • Human Evaluation (การประเมินโดยมนุษย์):
    ผู้เชี่ยวชาญ (Subject matter expert) ดู documents และ determine the relevance and accuracy of the outputs (ตัดสินความเกี่ยวข้องและความถูกต้องของผลลัพธ์) outputs
  • Noise Robustness (ความทนทานต่อสัญญาณรบกวน):
    ความสามารถของระบบ RAG ในการแยก noisy documents (documents ที่เกี่ยวข้องกับ query แต่ไม่มีข้อมูลที่เป็นประโยชน์) ออกจาก relevant ones
  • Negative Rejection (การปฏิเสธเชิงลบ):
    ความสามารถของระบบ RAG ในการ “ไม่ให้คำตอบ” เมื่อไม่มีข้อมูลที่เกี่ยวข้องกับ query ใน documents ใน knowledge base
  • Information Integration (การบูรณาการข้อมูล):
    ความสามารถของระบบในการ assimilate information (รวมข้อมูล) จาก multiple documents เพื่อตอบ query ได้อย่างครอบคลุม
  • Counterfactual Robustness (ความทนทานต่อข้อเท็จจริงที่ขัดแย้งกัน): ความสามารถของระบบ RAG ในการ address (จัดการ) และ reject known inaccuracies (ปฏิเสธความไม่ถูกต้องที่รู้แล้ว) ใน retrieved information

2️⃣ Frameworks (เฟรมเวิร์ก)

  • Frameworks (เฟรมเวิร์ก):
    เครื่องมือ (Tools) ที่ออกแบบมาเพื่อช่วยในการ evaluation โดยมี automation ของ evaluation process และ data generation
  • RAGAS (Retrieval Augmented Generation Assessment):
    framework ที่พัฒนาโดย Exploding Gradients ที่ assesses retrieval และ generation components ของ RAG systems โดยไม่ต้องใช้ human annotations มากนัก
  • Synthetic Test Dataset Generation (การสร้างชุดข้อมูลทดสอบสังเคราะห์):
    การใช้ models เช่น LLMs ในการ automatically generate ground truth data จาก knowledge base

Synthetic Data Generation in RAGAS (Source: A Taxonomy of Retrieval Augmented Generation)

  • LLM as a judge (LLM เป็นผู้ตัดสิน):
    การใช้ LLM ในการ evaluate a task (ประเมินงาน)
  • ARES (Automated RAG evaluation system):
    framework ที่พัฒนาโดย researchers ที่ Stanford University และ Databricks โดยใช้หลักการ LLM as a judge

3️⃣ Benchmarks (เกณฑ์มาตรฐาน)

  • Benchmarks (เกณฑ์มาตรฐาน):
    Standardised datasets และ evaluation metrics ที่ใช้ measure performance ของ RAG systems ซึ่งช่วย provide a common ground สำหรับ comparing different RAG approaches และ ensure consistency across evaluations โดย considering fixed tasks และ evaluation criteria
  • BEIR (Benchmarking Information Retrieval):
    A comprehensive heterogeneous benchmark ที่ based on 9 IR tasks และ 19 Question-Answer datasets

BEIR — 9 tasks and 18 (of 19) datasets (Source: BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models)

  • Retrieval Augmented Generation Benchmark (RGB):
    A benchmark ที่ focusses on 4 key abilities ของ RAG system — 1) Noise Robustness, 2) Negative Rejection, 3) Information Integration และ 4) Counterfactual Robustness

Four abilities required of RAG systems (Source: Benchmarking Large Language Models in Retrieval-Augmented Generation, Chen et al )

  • Multihop RAG:
    a benchmark ที่ contains queries ที่ requires reasoning across multiple documents โดยจะ queries involve document metadata, reflecting complex scenarios ที่ commonly found ใน real-world RAG applications
  • Comprehensive RAG (CRAG):
    a benchmark ที่ focusses on factual question answering และ simulates web และ Knowledge Graph (KG) search โดยจะ contains 8 types ของ queries across 5 domains

Pipeline Design (การออกแบบไปป์ไลน์)

1️⃣ Naive RAG (RAG แบบดั้งเดิม)

  • Naive RAG:
    RAG แบบพื้นฐานที่สุด ทำงานเป็นเส้นตรง (linear) ตามลำดับ: indexing -> retrieval -> augmentation -> generation
  • Retrieve-Read:
    A retriever ที่ดึงข้อมูล และ LLM อ่านข้อมูลนั้นเพื่อสร้างคำตอบ
  • RAG Failure Points (จุดบกพร่องของ RAG):
    - retriever ดึงข้อมูลไม่ครบ หรือดึงข้อมูลที่ไม่เกี่ยวข้อง
    - LLM ไม่สนใจ context ที่ให้มา หรือเลือกข้อมูลที่ไม่เกี่ยวข้องจาก context
  • Disjointed Context (บริบทที่ไม่ต่อเนื่อง):
    ข้อมูลมาจากหลายแหล่ง ทำให้การเชื่อมต่อระหว่าง chunks ไม่ราบรื่น
  • Over-reliance on Context (การพึ่งพาบริบทมากเกินไป):
    LLM ลืมความรู้เดิมของตัวเอง (parametric memory)

2️⃣ Advanced RAG (RAG ขั้นสูง)

  • Advanced RAG:
    RAG ที่ปรับปรุงขั้นตอนต่างๆ (pre-retrieval, retrieval, post-retrieval) เพื่อแก้ข้อจำกัดของ Naive RAG

Advanced RAG as Rewrite-Retrieve-Rerank-Read pattern (Source: A Taxonomy of Retrieval Augmented Generation)

  • Rewrite-Retrieve-Rerank-Read:
    เพิ่มขั้นตอนการ rewrite (เขียนใหม่) query และ rerank (จัดอันดับใหม่) ผลลัพธ์
  • Index Optimisation (การปรับปรุงดัชนีให้ดีที่สุด):
    จัดเตรียม knowledge base ให้ดีขึ้นสำหรับการ retrieval
  • Query Optimisation (การปรับปรุงคำถามให้ดีที่สุด):
    ปรับ query ให้เหมาะกับการ retrieval
  • Query Expansion (การขยายคำถาม):
    เพิ่มคำใน query เพื่อให้ได้ข้อมูลที่เกี่ยวข้องมากขึ้น
    - Multi-query expansion: สร้างหลาย variations ของ query
    - Sub-query expansion: แบ่ง query ซับซ้อน เป็น queries ย่อยๆ
    - Step back expansion: สร้าง query ที่เป็น conceptual query (query เชิงความคิดรวบยอด)
  • Query Transformation (การแปลงคำถาม):
    ใช้ query ที่ transformed (แปลงแล้ว) แทน query เดิม
  • Query Rewriting (การเขียนคำถามใหม่):
    เขียน query ใหม่จาก input ที่อาจไม่ใช่ query โดยตรง
  • Hypothetical document embedding, HyDE (การฝังเอกสารสมมติ): LLM สร้าง hypothetical answer (คำตอบสมมติ) แล้วใช้ answer นั้นในการค้นหา
  • Query Routing (การจัดเส้นทางคำถาม):
    ส่ง query ไปยัง workflow ที่เหมาะสม ตาม criteria ต่างๆ
  • Hybrid Retrieval (การดึงข้อมูลแบบผสมผสาน):
    ใช้หลาย retrieval methods ร่วมกัน (เช่น keyword-based search + semantic similarity)

Hybrid of sparse, dense and graph retrieval (Source: A Taxonomy of Retrieval Augmented Generation)

  • Iterative Retrieval (การดึงข้อมูลแบบวนซ้ำ):
    ดึงข้อมูลหลายครั้ง โดยใช้ generated response ในการดึงข้อมูลครั้งต่อไป
  • Recursive Retrieval (การดึงข้อมูลแบบเรียกซ้ำ):
    เหมือน Iterative Retrieval แต่มีการ transform query หลังจากการ generate แต่ละครั้ง
  • Adaptive Retrieval (การดึงข้อมูลแบบปรับตัว):
    LLM กำหนดเวลาและ content ที่เหมาะสมสำหรับการ retrieval
  • Contextual Compression (การบีบอัดบริบท):
    ลดความยาวของข้อมูลที่ retrieved โดยเลือกเฉพาะส่วนที่เกี่ยวข้อง
  • Explanation: Reducing the amount of retrieved information to only the relevant parts.
  • Reranking (การจัดอันดับใหม่):
    จัดอันดับ retrieved information เพื่อเลือก documents ที่เกี่ยวข้องที่สุด

3️⃣ Modular RAG (RAG แบบแยกส่วน)

  • Modular RAG:
    แบ่งโครงสร้าง RAG แบบเดิม (monolithic) ออกเป็น components ที่สามารถสลับเปลี่ยนได้ (interchangeable) เพื่อให้ปรับแต่งระบบได้ตาม use cases
  • Search Module (โมดูลการค้นหา):
    ค้นหาข้อมูลจากแหล่งต่างๆ
  • RAG-Fusion:
    ปรับปรุงระบบ search เดิม โดยใช้ multi-query approach
  • Memory Module (โมดูลหน่วยความจำ):
    ใช้ประโยชน์จาก “memory” ของ LLM (ความรู้ที่อยู่ใน parameters)
  • Routing (การจัดเส้นทาง):
    นำทาง query ผ่านแหล่งข้อมูลต่างๆ เพื่อเลือก pathway ที่เหมาะสม
  • Task Adapter (ตัวปรับงาน):
    ทำให้ RAG ปรับตัวเข้ากับ downstream tasks ต่างๆ ได้ เช่น summarisation, translation

Operations Stack (Operations Stack)

1️⃣ Critical Layers (เลเยอร์ที่สำคัญยิ่ง)

  • Critical Layers (เลเยอร์ที่สำคัญยิ่ง):
    องค์ประกอบพื้นฐานที่ RAG system ขาดไม่ได้ ถ้าขาดไป ระบบจะล้มเหลว
  • Data Layer (เลเยอร์ข้อมูล):
    สร้างและจัดเก็บ knowledge base โดย collect data จากแหล่งต่างๆ, transform ให้เป็น format ที่ใช้งานได้, และ store เพื่อ efficient retrieval
  • Model Layer (เลเยอร์โมเดล):
    - จัดเก็บ model library, training & fine-tuning components, และ inference optimisation components
    - ทำให้ LLM สามารถ generate คำตอบได้อย่างรวดเร็วและคุ้มค่า

Model Layer of the RAGOps stack (Source: A Taxonomy of Retrieval Augmented Generation)

  • Fully managed deployment (การ deployment แบบจัดการเต็มรูปแบบ):
    ผู้ให้บริการจัดการ infrastructure ทั้งหมด
  • Self-hosted deployment (การ deployment แบบ self-hosted):
    ผู้พัฒนาจัดการ infrastructure เอง บน private clouds หรือ on-premises
  • Local/edge deployment (การ deployment แบบ local/edge):
  • รัน model บน local hardware หรือ edge devices เพื่อ data privacy, reduced latency, และ offline functionality
  • Application Orchestration Layer (เลเยอร์ประสานงานแอปพลิเคชัน): จัดการ interactions ระหว่าง layers ต่างๆ, เป็น central coordinator ที่ enable communication ระหว่าง data, retrieval systems, generation models, และ services อื่นๆ

2️⃣ Essential Layers (เลเยอร์ที่จำเป็น)

  • Essential Layers (เลเยอร์ที่จำเป็น):
    เลเยอร์ที่ focus บน performance, reliability, และ safety ของระบบ ทำให้ระบบมีมาตรฐานและให้ value แก่ผู้ใช้
  • Prompt Layer (เลเยอร์พรอมต์):
    จัดการ augmentation และ LLM prompts ต่างๆ
  • Evaluation Layer (เลเยอร์การประเมินผล):
    จัดการ regular evaluation ของ retrieval accuracy, context relevance, faithfulness, และ answer relevance
  • Monitoring Layer (เลเยอร์การตรวจสอบ):
    - Continuous monitoring เพื่อ long-term health ของ RAG system
    - เข้าใจ system behaviour, identify points of failure, assess relevance & adequacy of information, และ track system metrics ต่างๆ
  • LLM Security & Privacy Layer (เลเยอร์ความปลอดภัยและความเป็นส่วนตัวของ LLM):
    - Ensure data privacy และ protection ด้วย strategies เช่น anonymisation, encryption, differential privacy, query validation & sanitisation, และ output filtering
    - Implement guardrails, access controls, monitoring, และ auditing
  • Caching Layer (เลเยอร์แคช):
    ลด cost และ latency โดย store frequently accessed data

RAGOps stack with critical and essential layers (Source: A Taxonomy of Retrieval Augmented Generation)

3️⃣ Enhancement Layers (เลเยอร์เสริม)

  • Enhancement Layer (เลเยอร์เสริม):
    เลเยอร์ที่ improving efficiency, scalability, และ usability ของระบบ เลือกใช้ตาม end requirements
  • Human-in-the-loop Layer (เลเยอร์มนุษย์ในวงจร):
    ให้ human judgment ใน use-cases ที่ต้องการ higher accuracy หรือ ethical considerations
  • Cost Optimisation Layer (เลเยอร์การปรับต้นทุนให้เหมาะสม):
    Manage resources efficiently สำหรับ large-scale systems
  • Explainability and Interpretability Layer (เลเยอร์ความสามารถในการอธิบายและตีความ):
    Provide transparency สำหรับ system decisions ใน domains ที่ requiring accountability
  • Collaboration and Experimentation Layer (เลเยอร์การทำงานร่วมกันและการทดลอง):
    Useful สำหรับ teams ที่ทำงานบน development และ experimentation

Emerging Patterns (รูปแบบที่เกิดขึ้นใหม่)

1️⃣ Knowledge Graphs (กราฟความรู้)

  • Knowledge Graph powered RAG (RAG ที่ขับเคลื่อนด้วยกราฟความรู้):
    ใช้โครงสร้าง knowledge graph เพื่อเพิ่ม contextual understanding, enhanced reasoning capabilities, และ improved explainability
  • Knowledge Graphs (กราฟความรู้):
    จัดระเบียบข้อมูลเป็น entities (วัตถุ, แนวคิด) และ relationships (ความสัมพันธ์) ในรูปแบบ structured manner
  • GraphRAG:
    An open-source framework ที่สร้าง knowledge graphs จาก source documents โดยอัตโนมัติ และใช้ knowledge graph นั้นในการ retrieval
  • Graph Communities (ชุมชนกราฟ):
    แบ่ง entities และ relationships เป็นกลุ่มๆ
  • Community Summaries (บทสรุปชุมชน):
    LLM generate summaries สำหรับ communities เพื่อ insights into topical structure และ semantics
  • Local Search (การค้นหาในท้องถิ่น):
    หา set ของ entities ที่ semantically-related กับ user input จาก knowledge graph
  • Global Search (การค้นหาระดับโลก):
    Similarity based search บน community summaries
  • Ontology (ออนโทโลยี):
    A formal representation of knowledge as a set of concepts ภายใน domain, และ relationships ระหว่าง concepts เหล่านั้น

2️⃣ Multimodal (มัลติโมดอล)

  • Multimodal RAG (RAG แบบมัลติโมดอล):
    ใช้ modalities อื่นๆ นอกเหนือจาก text (เช่น images, audio, video) ใน both retrieval และ generation
  • Modality (โมดอลลิตี้):
    Specific type ของ input data (เช่น text, image, video, audio)
  • Multimodal Embeddings (Embeddings แบบมัลติโมดอล):
    A unified vector representation ที่ encode multiple data types (เช่น text และ image embeddings combined)
  • CLIP (Contrastive Language-Image Pre-training):
    A model ที่ learns visual concepts จาก natural language supervision, ใช้สำหรับการ cross-modal retrieval และ generation
  • Contrastive Learning (การเรียนรู้แบบเปรียบเทียบ):
    A learning method ที่ align data across different modalities โดย bringing semantically similar data points closer ใน shared embedding space

3️⃣ Agentic (แบบ Agent)

  • Agentic RAG (RAG แบบ Agent):
    Leverage LLM based agents สำหรับ adapting RAG workflow to query types และ type ของ documents ใน knowledge base
  • Adaptive Frameworks (เฟรมเวิร์กแบบปรับตัว):
    Dynamic systems ที่ adjust retrieval และ generation strategies based บน evolving context และ data
  • Routing Agents (เอเจนต์การจัดเส้นทาง):
    Agents responsible สำหรับ direct user queries to the most appropriate sources หรือ sub-systems
  • Query Planning Agents (เอเจนต์การวางแผนคำถาม):
    Agents ที่ break down complex queries into sub-queries และ manage execution across different retrieval pipelines
  • Multiple Vectors per Document (หลายเวกเตอร์ต่อเอกสาร):
    A technique ที่ multiple vector representations are generated multiple vector representations สำหรับ each document เพื่อ capture different aspects ของ content

Technology Providers (ผู้ให้บริการเทคโนโลยี)

  • Model Access, Training & FineTuning (การเข้าถึงโมเดล, การฝึกฝน, และการปรับแต่ง):
    ผู้ให้บริการที่ให้ access เข้าถึง, train, และ fine-tune LLMs รวมถึง cloud providers และบริษัทที่พัฒนา LLMs เอง เช่น OpenAI, HuggingFace ซึ่งเป็นแหล่งของ models ที่ใช้ใน RAG systems
  • Data Loading (การโหลดข้อมูล):
    ผู้ให้บริการ tools และ services สำหรับ loading และ preparing data เพื่อใช้ใน RAG systems เช่น LlamaIndex, LangChain ซึ่งช่วยให้การเตรียมข้อมูลเป็นไปอย่างมีประสิทธิภาพ
  • Vector DB and Indexing (ฐานข้อมูลเวกเตอร์และการทำดัชนี):
    ผู้ให้บริการ vector databases และ indexing solutions ที่สำคัญสำหรับการ efficient retrieval ใน RAG systems เช่น Pinecone, Chroma ซึ่งเป็น infrastructure หลักสำหรับการค้นหาข้อมูลที่เกี่ยวข้อง
  • Application Framework (เฟรมเวิร์กแอปพลิเคชัน):
    Frameworks ที่ provide tools และ components สำหรับ building RAG applications เช่น LangChain, LlamaIndex, CrewAI (Agentic Orchestration), LangGraph (Agentic Orchestration) ซึ่งช่วยลดความซับซ้อนในการพัฒนาระบบ RAG
  • Prompt Engineering (วิศวกรรมพรอมต์):
    ผู้ให้บริการ tools และ platforms สำหรับ prompt engineering ที่ช่วย optimize LLM performance เช่น W&B (Weights & Biases), PromptLayer ซึ่ง prompt ที่ดีช่วยให้ LLM ทำงานได้ดีขึ้น
  • Deployment Frameworks (เฟรมเวิร์กการ Deployment):
    Frameworks ที่ facilitate deployment ของ LLMs และ RAG systems เช่น MLflow ซึ่งช่วยให้การนำ LLMs และ RAG ไปใช้งานจริงง่ายขึ้น
  • Deployment & Inferencing (การ Deployment และการอนุมาน):
    Cloud providers และ services ที่ offer deployment และ inference capabilities สำหรับ LLMs เช่น AWS, GCP, OpenAI API, Azure ซึ่งเป็น infrastructure สำหรับรัน LLMs และ RAG systems
  • Monitoring (การตรวจสอบ):
    ผู้ให้บริการ monitoring tools และ services ที่ช่วย ensure health และ performance ของ RAG systems เช่น HoneyHive, TruEra ซึ่งช่วยให้ระบบทำงานได้อย่างราบรื่นและมีประสิทธิภาพ
  • Proprietary LLMs/VLMs (LLMs/VLMs ที่เป็นกรรมสิทธิ์):
    รายชื่อ LLMs และ Vision Language Models (VLMs) ที่เป็นกรรมสิทธิ์ (closed-source) เช่น GPT series by OpenAI, Gemini series by Google ซึ่งเป็นแหล่งของ models ที่มี license กำกับ
  • Open Source LLMs (LLMs โอเพนซอร์ส):
    รายชื่อ LLMs ที่เป็น open source สามารถนำไปใช้และแก้ไขได้อย่างอิสระ เช่น Llama series by Meta ซึ่งเป็นทางเลือกสำหรับผู้ที่ต้องการความยืดหยุ่นและควบคุม
  • Small Language Models (แบบจำลองภาษาขนาดเล็ก):
    รายชื่อ LLMs ที่มีขนาดเล็กกว่า เหมาะกับการใช้งานที่ต้องการความรวดเร็วและประหยัดทรัพยากร เช่น Gemma series by Google AI ซึ่งเป็นทางเลือกสำหรับ edge devices และสภาพแวดล้อมที่มีทรัพยากรจำกัด
  • Managed RAG solutions (โซลูชัน RAG ที่มีการจัดการ):
    โซลูชัน RAG ที่มีการจัดการทั้งหมด ช่วยให้การ setup และ management ของ RAG systems ง่ายขึ้น เช่น OpenAI File Search, Azure AI File Search ซึ่งเหมาะสำหรับผู้ที่ต้องการความสะดวกและไม่ต้องดูแล infrastructure เอง
  • Knowledge Graph and Ontology (กราฟความรู้และ Ontology):
    ผู้ให้บริการ knowledge graph databases และ ontology management tools เช่น Neo4j ซึ่งเป็น infrastructure สำหรับสร้างและจัดการ knowledge graphs
  • Security and Privacy (ความปลอดภัยและความเป็นส่วนตัว):
    ผู้ให้บริการ security และ privacy solutions สำหรับ AI systems รวมถึง RAG เช่น Hazy, Duality, BigID ซึ่งช่วย protect ข้อมูลและ ensure compliance
  • Synthetic Data (ข้อมูลสังเคราะห์):
    ผู้ให้บริการ synthetic data generation tools ที่สามารถใช้ augment training datasets และ improve model performance เช่น Mostly AI, Tonic.ai, Synthesis AI ซึ่งช่วยเพิ่มปริมาณและคุณภาพของข้อมูลสำหรับ training
  • Others (อื่นๆ):
    หมวดหมู่สำหรับ technologies และ services อื่นๆ ที่เกี่ยวข้องกับ RAG เช่น Cohere reranker, Unstructured.io

Applied RAG (RAG ที่นำไปใช้)

1️⃣ Other RAG Patterns (รูปแบบ RAG อื่นๆ)

  • Corrective RAG (RAG ที่แก้ไขได้):
    ดึงข้อมูล real-time เพื่อ check factual accuracy ของ LLM generated answer (ใช้ในการ fact-checking, medical & legal domains)
  • Contrastive RAG (RAG แบบเปรียบเทียบ):
    ใช้ contrastive learning เพื่อ enhance retrieval process โดยแยก relevant และ irrelevant documents
  • Selective RAG (RAG แบบคัดเลือก):
    Optimise retrieval phase โดย determine when it is beneficial to retrieve external information (ใช้ใน context ที่ retrieval อาจไม่มี value)
  • RAG with Active Learning (RAG กับการเรียนรู้เชิงรุก):
    ใช้ user feedback เพื่อ fine-tune หรือ adapt retrieval process over time (ใช้ใน continuous improvement systems เช่น recommendation engines)
  • Personalised RAG (RAG ส่วนบุคคล):
    ใช้ user preferences, behaviour, และ historical interactions เพื่อ personalise retrieval process (ใช้ใน personalization-heavy domains เช่น recommendation engines, customer service)
  • Self-RAG:
    Adaptive retrieval mechanism ที่ selectively decides when to retrieve knowledge based บน query’s context
  • RAFT (Retrieval-Augmented Fine-Tuning):
    Combine retrieval mechanisms with traditional fine-tuning techniques
  • RAPTOR (Recursive Abstractive Processing for Tree-Organised Retrieval):
    สร้าง recursive, tree-like structure จาก documents เพื่อ improve context-aware information retrieval

2️⃣ Application Areas (ขอบเขตการใช้งาน)

  • Search Engine (เครื่องมือค้นหา):
    ใช้ RAG เพื่อ present coherent text ใน natural language with source citation

Prominent Search Engines using RAG (Source: Google.com, Bing.com, Perplexity.ai, OpenAI.com)

  • Personalised Marketing Content Generation (การสร้างเนื้อหาทางการตลาดส่วนบุคคล):
    Content can be personalised to readers, incorporate real-time trends, and be contextually appropriate
  • Personalised Learning Plans (แผนการเรียนรู้ส่วนบุคคล):
    สร้าง personalised learning paths based บน past trends และ automated evaluation and feedback
  • Real-time Event Commentary (ความคิดเห็นเกี่ยวกับเหตุการณ์แบบเรียลไทม์):
    Connect to real-time updates/data via APIs และ pass information to LLM to create virtual commentator
  • Conversational agents (ตัวแทนสนทนา):
    Customise LLMs to product/service manuals, domain knowledge, guidelines, etc. และ serve as support agents
  • Document Question Answering Systems (ระบบตอบคำถามเกี่ยวกับเอกสาร):
    Answer all questions about the organisation with access to proprietary documents
  • Virtual Assistants (ผู้ช่วยเสมือน):
    Enhance user’s experience with more context บน user behaviour using RAG

3️⃣ Applied RAG Challenges (ความท้าทายในการใช้ RAG)

  • Relevance Mismatch (ความไม่ตรงกันของความเกี่ยวข้อง):
    Difficulty retrieving most relevant documents due to suboptimal ranking
  • Over-Retrieval (การดึงข้อมูลมากเกินไป):
    Retrieving too many documents, leading to unnecessary noise
  • Sparse vs Dense Retrieval Trade-off (การแลกเปลี่ยนระหว่างการดึงข้อมูลแบบเบาบางและหนาแน่น):
    Balancing between sparse retrieval และ dense retrieval เพื่อ maximise relevance
  • Document Question Answering Systems (ระบบตอบคำถามเกี่ยวกับเอกสาร):
    Delays due to retrieval จาก large knowledge bases
  • Latency (ความหน่วง): Retrieval จาก large หรือ distributed knowledge bases can introduce significant delays affecting real-time applications
  • Cost of Storage (ค่าใช้จ่ายในการจัดเก็บ):
    Maintaining massive vector databases can be expensive
  • Narrow Retrieval Focus (การโฟกัสการดึงข้อมูลที่แคบ):
    Difficulty retrieving diverse perspectives
  • Bias in Retrieval (อคติในการดึงข้อมูล):
    Biases ใน retrieval results based บน structure ของ data
  • Context Loss in Long Queries (การสูญเสียบริบทในคำถามยาว):
    Loss of context when handling long, multi-turn queries
  • Incoherent Summarisation (การสรุปที่ไม่สอดคล้องกัน):
    Generating inconsistent summaries จาก multiple documents
  • Over-Generation (การสร้างข้อมูลมากเกินไป):
    Generating overly verbose responses
  • Inconsistent Modal Alignment (การจัดแนว Modal ที่ไม่สอดคล้องกัน): Challenges integrating multimodal data
  • Data Silos (ไซโลข้อมูล):
    Knowledge is fragmented across multiple sources
  • Processing Large-Scale Data (การประมวลผลข้อมูลขนาดใหญ่): Difficulty maintaining high throughput as data grows
  • Multi-Agent Coordination (การประสานงานหลาย Agent):
    Complex coordination among multiple agents
  • Inefficient Query Routing (การจัดเส้นทางคำถามที่ไม่มีประสิทธิภาพ): Routing queries to wrong sources
  • Data Poisoning Attacks (การโจมตีด้วยการวางยาข้อมูล):
    External sources feed biased data into the generation pipeline
  • Adversarial Attacks (การโจมตีแบบ Adversarial):
    Attackers influence retrieval or generation results
  • Knowledge Base Updating (การอัปเดตฐานความรู้): Maintaining an up-to-date knowledge base
  • Memory Retention (การเก็บรักษาความทรงจำ):
    Ensuring system can store and retrieve long-term memory

หวังว่าบทความนี้จะช่วยให้คุณเข้าใจ RAG ได้อย่างลึกซึ้งและนำไปประยุกต์ใช้ได้อย่างมีประสิทธิภาพนะครับ แล้วพบกันใหม่ในบทความต่อไป 🚀


Data Science Explore the world of data science with Donato_Story

Dashboard Discover the power of data visualization with Donato_Story

Donato_Journey Join me on my journey (Thai version)

Course_Review Discover the training courses with Donato_Story (Thai version)

Let’s Connect!

Your thoughts and feedback are invaluable. Feel free to share them in the comments or connect with me on

Originally published on Medium

Related