В прошлой статье серии мы сжимали эмбеддинги и замеряли, насколько изменится выдача относительно полного fp32-поиска. Там это был правильный вопрос: ломает ли квантизация уже существующий retrieval.Но у такого замера есть неприятное слепое пятно. Можно аккуратно сохранить 99% выдачи…
Sharp NEC Displays (UN462A R1.300 and prior to it, UN462VA R1.300 and prior to it, UN492S R1.300 and prior to it, UN492VS R1.300 and prior to it, UN552A R1.300 and prior to it, UN552S R1.300 and prior to it, UN552VS R1.300 and prior to it, UN552 R1.300 and prior to it, UN552V R1.300 and prior to it, UX552S R1.300 and prior to it, UN552 R1.300 and prior to it, V864Q R2.000 and prior to it, C861Q R2.000 and prior to it, P754Q R2.000 and prior to it, V754Q R2.000 and prior to it, C751Q R2.000 and prior to it,
Sharp NEC Displays (UN462A R1.300 and prior to it, UN462VA R1.300 and prior to it, UN492S R1.300 and prior to it, UN492VS R1.300 and prior to it, UN552A R1.300 and prior to it, UN552S R1.300 and prior to it, UN552VS R1.300 and prior to it, UN552 R1.300 and prior to it, UN552V R1.300 and prior to it, UX552S R1.300 and prior to it, UN552 R1.300 and prior to it, V864Q R2.000 and prior to it, C861Q R2.000 and prior to it, P754Q R2.000 and prior to it, V754Q R2.000 and prior to it, C751Q R2.000 and prior to it,
Имея некоторый опыт в построении классических ML и CV-проектов, я решил разобраться в NLP (Natural Language Processing) и собрать свою RAG-систему без использования сторонних RAG-фреймворков (LangChain, LlamaIndex и т.п.). Моя цель - понять, как на самом деле работает RAG под капотом: от разбиения текста до генерации ответа. Проблема, которую решает RAG Готовые LLM модели имеют несколько недостатков: Читать далее