<img width="945" alt="image" src="https://github.com/InfiniteAICreations/ViperInsight/assets/16201837/a00dce80-4add-4516-b73c-e0defc1434c7"> <img width="1041" alt="image" src="https://github.com/InfiniteAICreations/ViperInsight/assets/16201837/0f6f63b2-d5ea-46ae-9b92-20b7844b0c1a"> The [Ferret-UI][1] is a good reference. Combining the low resolution and the sub-image makes sense in some situations. However, we lack the datasets, which are important to the LLMs fine-tuning or evaluating the result of inference. [1]: https://arxiv.org/pdf/2404.05719.pdf
The Ferret-UI is a good reference. Combining the low resolution and the sub-image makes sense in some situations.
However, we lack the datasets, which are important to the LLMs fine-tuning or evaluating the result of inference.