Agreed, this like asking a chainsaw to carve a wooden spoon. Impressive it can, but definitely not the right tech to scale.
LLM needs to setup an image classifier to use as a tool call.
Agreed, this like asking a chainsaw to carve a wooden spoon. Impressive it can, but definitely not the right tech to scale.
LLM needs to setup an image classifier to use as a tool call.
Building a dataset is expensive, manual annotation is expensive. Datasets don't exist in every niche.
I remember around 2013-15 people were scoffing at uses of deep learning CNNs for various things, because why don't you just use an SVM on HOG features? Or face detection is solved, just use Viola-Jones.
What if you give the benefit of doubt and assume the author knows about alternatives and uses VLMs for their strengths? They use it to auto-annotate training data for regular deep learning models.
Now maybe, but the gap is closing.