Skip to content
,

Photos, screenshots, PDFs, and multimodal basics

8 min read
Photos, screenshots, PDFs, and multimodal basics, with the official product logo. Editorial illustration for Analytics Made Simple.

You point your phone at a whiteboard after a meeting, ask Gemini what the action items were, and get a tidy list. That list can be brilliant. It can also invent an owner who never volunteered. Multimodal features feel like magic because they collapse camera, files, and language into one request. Magic is exactly when verification habits tend to fall asleep.

This is Part 5 of Learn Gemini from scratch. You will learn practical patterns for photos, screenshots, PDFs, and other files in the Gemini app and related Google surfaces, plus a safe upload ladder so curiosity does not become a data incident.

What multimodal means in plain English for Gemini

  • What multimodal means in plain English for Gemini
  • Good jobs for photos, screenshots, PDFs, and tables
  • Prompt patterns that force structure instead of vibes
  • A four-rung safe upload ladder
  • How to verify visual and document answers
  • When multimodal is the wrong tool

Multimodal without jargon

Multimodal means the product can take more than plain typed text as input: images, screenshots, PDFs, spreadsheets, audio in some modes, and combinations of those with instructions. Output is still mostly text (and sometimes generated images or other media, depending on plan and features). The model does not “see” the way a careful human auditor sees. It produces a plausible description and answer. Plausible is not the same as correct.

InputAsk it toDo not treat as
PhotosRead a whiteboard or product shotA witness
ScreenshotsExplain an error dialog or UIA bug ticket you can close
PDFs / docsMap structure and extract a sectionGround truth
SheetsExplain columns and suggest cleanupA signed-off model

Photos: whiteboards, products, real-world scenes

Good first jobs

  • Transcribe a whiteboard into bullets (then you clean names)
  • Explain a cable setup or appliance label in plain language
  • Describe a slide photo when you cannot edit the original file
  • Extract a packing list from a photo of a handwritten note
Here is a photo of a whiteboard.
1) Transcribe visible text as faithfully as you can
2) List action items as owner / task / due date if present
3) Mark anything unreadable as [UNCLEAR]
4) Do not invent owners or dates

The [UNCLEAR] instruction matters. Without it, models fill gaps with confident fiction.

Photo mistakes

  • Capturing badges, faces of non-consenting people, or desk papers with secrets
  • Asking for legal interpretation of a photographed contract
  • Using a blurry photo of a chart as financial ground truth

Screenshots: errors, UI confusion, charts

Screenshots are the everyday superpower. Error dialogs, confusing settings pages, and “what does this button mean?” questions are fair game when the screen does not contain secrets.

Screenshot attached.
Explain what the UI is asking me to do in 5 steps.
List risks if I click the primary button.
Ask me 3 questions if the screenshot is missing context.
Do not invent account-specific data.

For charts, ask for reading help, not authority: “Describe what this chart appears to show and what I should verify in the underlying data.” Then open the sheet.

PDFs and long documents

Upload or attach a PDF when you need structure: sections, definitions, open questions, comparison tables. Prefer documents you created or public materials for practice.

Ask forWhyVerify by
Section mapOrients you fastSkimming the PDF headings
Glossary of terms usedShared languageChecking one term in context
Open questions listShows gapsMarking which questions are real
Quote + paraphrase pairsGrounds claimsSearching the PDF for the quote
Using the attached PDF:
- Outline sections in under 12 bullets
- Extract 5 claims that sound quantitative
- For each claim, quote the nearest supporting sentence or say NOT FOUND
- List 5 follow-up questions for a human expert

If the model cannot quote support, treat the claim as untrusted. Smooth paraphrases without quotes are where hallucinations hide.

Tables and sheet-like files

When you attach CSV or sheet data (where supported), ask for explanations and checks, not silent rewrites of your system of record.

  • “Which columns look like IDs vs measures?”
  • “Propose validation rules for missing values.”
  • “Write a plain-English description of grain: what does one row mean?”

Then implement rules in Sheets or your warehouse tools. AI suggestions are design help.

Safe upload ladder

RungExamplesDefault action
NeverSSN, full card numbers, passwords, raw medical IDsDo not upload
High riskCustomer lists, HR files, unredacted contractsUse approved work tools only, or do not use AI
CarefulInternal decks with names; partially sensitive screenshotsRedact first; prefer Workspace under policy
SaferPublic PDFs, your own non-sensitive photos, synthetic samplesOK for learning and many drafts

When in doubt, describe the document instead of uploading it: “I have a 12-page vendor MSA with unlimited liability language. What clauses should a human lawyer review?” You still need the lawyer.

Verification habits for multimodal answers

  1. Require quotes or [UNCLEAR] labels for document and photo transcription jobs.
  2. Check at least one surprising claim against the source file.
  3. For numbers, re-type them from the source, do not trust the model’s table copy blindly.
  4. For action items from whiteboards, confirm owners in the next human conversation.
  5. Keep the source file linked next to the AI summary in your notes.

When multimodal is the wrong tool

  • You need a certified accessible transcript with legal standing
  • The image is evidence in a dispute and chain of custody matters
  • The PDF is confidential and your AI surface is not approved
  • You are trying to OCR thousands of pages (use dedicated pipelines)
  • You want pixel-perfect data extraction for production systems without human review

Take a photo of a public sign

  1. Take a photo of a public sign or a page of a public PDF on screen (not sensitive).
  2. Run the transcription prompt with [UNCLEAR] rules.
  3. Compare to what you can read yourself; note one error if any.
  4. Summarize a short public PDF with quote + NOT FOUND rules.
  5. Write your personal upload rule in one sentence and keep it visible.

Worked example: screenshot of a confusing settings page

You are enabling two-factor authentication and the screen has three toggles with unclear labels. Capture a screenshot that excludes unrelated personal data. Then:

Screenshot of a settings page attached.
Goal: I want stronger login security without locking myself out.
1) Name each visible toggle in plain English
2) Recommend a safe order of operations in 5 steps
3) List what I should write down before I change anything
4) Call out any toggle that looks destructive
Ask questions if the screenshot is incomplete.

Compare the answer to the vendor’s official help article for that product. AI is a navigator. The vendor docs remain the authority for security settings.

Whiteboards and meetings

Photographing a whiteboard is faster than typing. It is also how you capture personal phone numbers and side jokes that should not enter a cloud chat. Before you shoot, erase or crop sensitive bits. After Gemini drafts action items, paste them into the meeting notes doc and assign owners in the room or the thread. An AI list with no owners is decoration.

If your company uses Meet note features, the same rule applies: AI notes are a draft. Decisions need a human stamp.

PDF research workflow that scales to real work

  1. State the question before you upload (“I need termination notice periods only”).
  2. Ask for a section map first.
  3. Ask for quote-backed answers to your question.
  4. Build a small table: claim | quote | page/section | confidence.
  5. Anything without a quote stays out of your final memo.

This is slower than “summarize this PDF” and much harder to get wrong in public. Speed without provenance is how confident errors spread in Slack.

Images you generate vs images you upload

Some Gemini plans include image generation or editing features. Generated images have their own brand and policy issues (likeness, trademarks, misleading realism). Uploaded images have privacy issues. Do not mix those risk models. For teaching diagrams on AMS we often build charts in code for exact labels; for personal brainstorming, generated images can be fine if you are not impersonating a real person or faking evidence.

Accessibility and fairness notes

Automated descriptions of images can miss critical content or misread text. If you rely on AI descriptions for access, pair them with human checks for anything important. Also remember bias: models can misidentify people, culture-specific items, or medical images. Those domains need specialized tools and professionals, not a general chat upload.

Voice and live-style features (when you have them)

Some Gemini experiences include voice conversation or camera-assisted live help. Those modes are fantastic for cooking timers, travel confusion, and hands-busy troubleshooting. They are also easy to leave running in a room where other people are discussing private topics. Treat the microphone like an open meeting guest: introduce the tool, mute when the conversation is not for it, and never point a live camera at badges, mail, or medical paperwork.

If voice transcripts appear in history, apply the same deletion and sensitivity rules as text chats. A spoken customer name is still a customer name.

Building a personal multimodal kit

  • A redaction habit: crop first, upload second
  • A notes template with Source file / AI summary / Verified claims columns
  • A folder of public sample PDFs for practice so you never need real customer files to learn
  • A team agreement on whether whiteboard photos are allowed after meetings
  • A default prompt snippet with [UNCLEAR] and quote rules saved in your notes app

Skills compound. The first week feels slow. The fourth week you will refuse bad uploads automatically and still move faster than peers who paste everything.

Multimodal is powerful and still probabilistic

  • Multimodal is powerful and still probabilistic.
  • Structure prompts; demand quotes and uncertainty labels.
  • Follow the upload ladder; never is a real category.
  • Next: Part 6, privacy and Workspace admin basics for normal users.

Sources