Skip to content
,
Grok · Part 5

Voice and everyday multimodal capabilities

10 min read
Voice and everyday multimodal capabilities, with the official product logo. Editorial illustration for Analytics Made Simple.

Voice and everyday multimodal capabilities

Modern AI tools can read a photo, a screenshot or a sketch and turn it into typed text, a data table or working code. That saves you the slow job of retyping things by hand. This post shows what that looks like for receipts, whiteboards and rough drawings, and where it still goes wrong.

Picture a photo on your phone of a whiteboard from a two-hour brainstorming meeting. It is covered in overlapping arrows, half-finished abbreviations and sticky notes that are barely holding on. Turning that into clean notes used to mean forty-five minutes of squinting and typing, and guessing whether a coworker wrote “sync API” or “sink apples.” It is the kind of chore that drains the energy out of a project.

The word for the fix is multimodal, which just means a tool can handle more than one kind of input, such as text and pictures, at the same time. It can look at an image, read the words in it, and write something new in response. You can hand it a photo of a whiteboard, a crumpled receipt or a rough app sketch on a napkin, and get back something you can use. The goal is not to be impressed by a computer describing a cat photo. It is to move quickly from a messy real-world object to something a computer can work with.

Reading receipts, whiteboards and sketches

Start with what these tools do to a receipt, a whiteboard or a mockup of an app screen. The model does more than recognize that a receipt is a receipt. It also understands where things sit on the page. It knows that the number to the right of “Total” is the final price, even when a coffee stain covers the decimal point.

Older computer vision programs relied on hand-written rules, such as find a dollar sign and read the numbers next to it. If the receipt was tilted or crumpled, the rules broke. Modern models do not use those rules. Instead, the image is cut into a grid of small patches, and each patch is turned into numbers the model can process next to your written prompt (these number lists are called embeddings). Because the model looks at the picture and your instructions together, it can connect what you asked with what it sees.

Now think about an app mockup. You draw a square, put a smaller rectangle inside it, and write “Submit” on the rectangle. A good model sees a button inside a box, and it understands what you meant by the drawing, not only the shapes. You can then ask for the matching HTML and CSS, the two languages that describe how a web page looks. This works well because the model learned from millions of pairs of screen designs and the code behind them, so it can match the layout of your sketch to the structure a browser needs.

The same skill works on architecture diagrams. Upload a photo of a cloud system sketch and the model can pick out the databases, the load balancers (which spread traffic across servers) and the API gateways (which control how programs connect). It understands that a line between two boxes usually means data moving from one to the other. A flat picture becomes something you can ask questions about.

The diagram below shows how the system takes different kinds of images and turns them into results you can use.

Turning messy photos into structured data

Optical character recognition (OCR, software that reads printed characters from an image) has existed for decades. It works well on clean printed text on flat paper. It struggles with curved pages, bad lighting and messy handwriting.

Modern tools work differently, because they do not only match letter shapes. They predict which characters ought to be there based on the whole image. If a thumb covers part of a word, the model can guess the missing letters from the sentence around it, much as you would.

You can also ask for the answer in a set shape. Instead of “give me all the text in this image,” you ask for a JSON object with specific fields (JSON is a plain text format that spreadsheets and programs can read). That turns a typing job into a step you can feed straight into another system. Suppose you photograph a restaurant menu and use a prompt like this:

Extract all the appetizers from this menu image. 
Return the result as a JSON array of objects. 
Each object should have "name" (string), "price" (number), and "is_vegetarian" (boolean).
Infer the vegetarian status from the item description. Do not include markdown formatting in the response.

In one pass, the model reads the words, organizes them into fields, and makes a judgment about which dishes are vegetarian based on the ingredients. For anyone who does data entry or turns paper into digital records, that is a big time saver.

The same idea applies to printed tables. If you have a photo of a quarterly earnings report, you can ask for the rows and columns as a clean CSV (a simple text file that opens in a spreadsheet). A slight angle in the photo is not a problem, because the model adjusts for the tilt. In many everyday cases, that removes the need for someone to retype the numbers.

The table below pairs common photos with the result you can ask for and the reason it helps at work.

Visual inputTypical output taskReal-world value
Coffee shop receiptJSON payload for expense softwareAutomated accounting and fast reimbursement
Whiteboard flowchartMermaid.js diagram codeInstant documentation without drawing tools
Hand-drawn UI sketchReact component with Tailwind CSSRapid prototyping and faster work on the visible part of a website
Vintage newspaper clippingClean transcribed markdown textArchival digitization and historical research
Error screen photoTroubleshooting steps and CLI commandsFaster IT support and less downtime
Printed data tableFormatted CSV stringFewer typing mistakes from manual data entry

Making original images with Flux models

Everyday multimodal tools also make pictures, not only read them. Image generation has improved quickly, and Flux models are a clear step up in quality and in following your instructions closely.

A Flux model works by a method called diffusion. It starts with a canvas of random visual static, then reads your prompt and removes the static in small steps until a picture appears that matches your description. You have real control over the result. You can specify the lighting, the camera angle, the type of film and where each subject sits, so you are no longer rolling dice and hoping.

Great results depend on being very specific. A prompt like “a dog in a park” gets a bland picture that looks obviously machine-made. A prompt like “A golden retriever catching a red frisbee mid air in a sunny park, shot on 35mm film, low angle, shallow depth of field, warm morning light” gets something that looks like a professional photograph. Think of yourself as a film director telling the crew exactly what to shoot.

People use this for blog images, custom icons for an app, and early product design ideas before anyone builds a physical prototype. When you ask for a graphic or screen design, name the color palette and the artistic style. Here is an example:

A flat vector illustration of a smartphone displaying a dashboard. 
Use a monochromatic blue color palette. 
The style should be minimalist, corporate, and suitable for a SaaS landing page. 
Clean lines, no gradients, pure white background.

Limits like these keep the model from falling back on its default look, which is often too detailed or dreamlike. The aim is a useful image, not abstract art, and a useful image can go straight into Figma or a slide deck without more editing.

Checking images against current information from the web

Older models had a serious weakness. They only knew what they had read up to a fixed date, called the knowledge cutoff. Ask one to draw the newest phone and it would guess from old information. Show it a photo of a breaking news event and it would not know what was going on.

Connecting the model to live web search fixes much of that. It can look at an image, recognize a landmark or a product, search the web for the latest news or specifications, and answer in a way that matches both the picture and today’s facts. Researchers call the general idea retrieval-augmented generation, meaning the model looks something up before it answers. Here it is applied to pictures.

Say you upload a photo of a concert stage and ask who is playing tonight. The tool can read the banners in the photo, check the venue’s current schedule online, and give you the right name. The image gives it a starting point, and the search supplies the facts.

This is especially handy for tech support. If you photograph a circuit board and point at a burnt part, the model can identify the board, search the maker’s latest documentation, and tell you which resistor to replace and where to order one. It connects recognizing a picture with solving a live problem.

Five mistakes that trip people up

Multimodal AI is powerful, but it is easy to stumble if you do not know its limits. These are the problems developers and everyday users hit most often.

  • Low resolution inputs. The model cannot pull text from a blurry or out-of-focus photo, and poor input gives poor output. Make sure your photo is well lit and that you can read it yourself first. If you cannot read it, the AI probably cannot either.
  • Trusting made-up text in generated images. Image generators are notoriously bad at spelling. Ask for a sign that says “Welcome to Chicago” and you might get “Welcom to Chicgo” or a jumble of strange symbols. Always check any words in a generated image, and consider adding the text afterward in an image editor.
  • Overloading the prompt. When you extract data, do not ask for fifty fields at once. Break the job into steps by asking for the basic structure first, then running a second prompt for the details. Combining a hard reading task with a hard thinking task often leads to skipped fields.
  • Ignoring privacy limits. Do not upload photos with sensitive personal information, passwords, financial documents or confidential company data, unless you are on a business plan with clear, guaranteed rules about how long data is kept. Public consumer tools often use what you give them to train future models.
  • Assuming perfect spatial awareness. Models handle simple screen mockups well, but they can struggle with crowded diagrams where lines cross a lot. Always check the relationships in the code it gives you.

Try it: turn a sketch into a web page

Now put this into practice with a small test. You will see how well a model turns a drawing into working code.

Grab a pen and paper and draw a rough layout for a simple profile card. Include a circle for a photo, a block for a name, a smaller block for a short bio, and a “Follow” button. It does not need to be neat, and a messy sketch is fine.

Take a photo of your drawing and open your favorite multimodal AI tool. Upload the image and use this exact prompt:

Analyze this hand-drawn UI sketch. 
Write the HTML and inline CSS required to recreate this layout as a functional web component. 
Make the design look modern, with rounded corners, a soft drop shadow, and a minimalist aesthetic.
Center the card on the page. Only output the HTML and CSS code block.

Look over what comes back, then save the code in a file ending in .html and open it in your browser. You will see how quickly a sketch can become a working page without any design software in between.

Quick recap: six habits for working with images and AI

  • Use multimodal models to pull clean text and structured formats like JSON out of messy photos of receipts, documents and whiteboards.
  • Take advantage of the model’s sense of layout to turn screen mockups straight into working front-end code.
  • Write very specific prompts when you use Flux models, so you control the lighting, composition and style.
  • Combine images with live web search, so the answers rest on current, factual information.
  • Always check the words inside generated images, because models often get the spelling wrong.
  • Give the model clear, well-lit photos to get the most accurate data out of them.

Sources

  • OpenAI documentation on Vision API best practices: https://platform.openai.com/docs/guides/vision
  • Anthropic research on multimodal capabilities: https://www.anthropic.com/news/multimodal-claude
  • Google Cloud guide to structured output from Gemini: https://cloud.google.com/vertex-ai/generative-ai/docs/multimodal/overview
Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: