vision in customer support: the ai that reads photos and payment receipts.
heysteff is an AI platform for customer support and sales across WhatsApp, Instagram, Messenger, Gmail and Shopify; Steff is the AI agent that runs it. A customer sending a photo instead of typing is one of the most common things in messaging — and an agent that can't "see" that image loses the most important context in the conversation.
◆ key idea
When the AI is not sure what it sees in an image, it doesn't make up an answer: it escalates to a human.
Case 1: the customer sends a photo of the product
"Will this color look like this on me?", "it arrived with this mark, is it damaged?", "is this size the M or the L?" — questions a customer resolves much faster by sending a photo than by describing them in words. An agent with vision can analyze that image within the same conversation, without asking the customer to upload it to another system or wait for a human to review it. That speeds up resolutions that would otherwise stall waiting for someone on the team to be available.
Case 2: payment receipt verified in the conversation
With bank transfers or payments outside an automated checkout — very common across several LATAM countries — the customer usually sends a screenshot of the receipt over WhatsApp. An agent with vision can read the relevant data from that image — amount, date, reference — as part of the payment confirmation flow, instead of someone on the team having to open the image manually and match it against the order. It's one of the most repetitive, error-prone tasks in messaging support, and vision helps it happen inside the same conversation.
Case 3: labels and manuals read in context
When a customer asks about a product by showing its label — ingredients, care instructions, batch number — or a page from a manual, the agent can read that text inside the image and reply with that specific information, instead of a generic catalog answer that may not match the exact version of the product the customer is holding.
How it works, at a high level: multimodal models
Computer vision in a conversational agent works through multimodal language models: the same kind of model that processes text can also receive an image and "describe" or reason about its content as part of the same message. It's not a separate image-recognition system duct-taped to the chatbot — it's a single capability of the AI model the agent uses depending on what the customer sends. In heysteff, this capability is enabled by the model chosen for the workspace; not all available models have vision of the same quality, so it's worth weighing that criterion when choosing the AI model for your agent.
The honest limits: when in doubt, it escalates
No vision model is perfect. A blurry or poorly lit photo, or a receipt with partially cut-off data, can produce a wrong reading if the agent "guessed" instead of admitting the doubt. That's why the expected behavior — and the one we configure in Steff — is not to force an answer when the image isn't clear: it's to say explicitly that it couldn't be read with confidence and escalate to a human or ask for a better photo, instead of risking a payment confirmation or a product answer based on a doubtful reading.
That honesty about limits is part of the same principle that applies to the whole agent: Steff would rather admit it doesn't know than invent a plausible but wrong answer, especially in something as sensitive as confirming a payment.
It also means computer vision doesn't replace human judgment in high-stakes decisions. A receipt that looks valid but has subtle inconsistencies, or a product photo that raises a warranty dispute, are cases where the AI's value lies in doing the fast first pass — reading, sorting, summarizing what it sees — and leaving the final call to a person when the amount or the situation warrants it. That combination, rather than full unsupervised automation, is what delivers reliable results in practice.
Related
◆ next step
Show us a real receipt or product photo and we'll show you how Steff responds.