I’m consulting for company that is building a solution to connect and ingest data from company websites and documents. We have found that low cost LLM models are often good enough to answer questions accurately about the data. It’s the context of the data (categories, purpose, origin, age) that affects the accuracy of the answers. Some of the challenging issues are determing the source of truth from similar data.
We are curious how others are approaching this problem, i.e. a connecting to a mixture of structured and free form text. Also, if Cdata has a solution or recommendation for this.

