Contributor Resources

Contributor Journey

From application to tasking: every stage of the path, and how to grow from here.

A

Getting started and setup

Complete the following steps to get fully set up:

  1. Set up the Workada Timer
    • Click "Set up" on the "Setup Workada Timer" card on your dashboard.
    • Click "Download app" and install the timer on your computer.
    • Open the app and click "Connect app" to link it to your account.
    • When you see "You're all set," click "Continue."
    Workada Timer app showing 00:00 with Start and Stop buttons and recent sessions
    Workada Timer app
    Workada Timer overlay mini-widget showing 00:16 with play and stop buttons
    Overlay mini-timer
  2. Set up payments
    • Go to the Payments tab on the left side of your dashboard.
    • Follow the prompts to connect your Stripe account.
    • Make sure this is done before your first payment cycle.
  3. Accept Terms of Use
    • Read and accept Workada's Terms of Use.
  4. Complete identity verification

    This takes about 2–5 minutes. Here's what you'll need to do:

    • Upload a valid, non-expired government-issued photo ID: driver's license, passport, or state/national ID. Have both the front and back ready to photograph.
    • Confirm your basic personal information (full name, address, date of birth): this is often extracted automatically from your ID.
    • Complete a selfie/liveness check: you'll take a live selfie and may be asked to turn your head or look at different points on the screen. This confirms you're physically present and that your face matches your ID.

    Most verifications are approved within minutes. Some cases may require manual review and take a little longer.

  5. Complete training

    Training is self-paced and completed on the Workada dashboard before you can access any live tasks. It's made up of multiple sections: each one builds on the last, with practice rounds followed by a short knowledge check. You need to pass each check to move forward.

    Here's what you'll cover:

    • What the work involves and what a good submission looks like
    • How to spot AI artifacts: the subtle mistakes models make in hands, faces, text, and backgrounds
    • The rating scale and how to use it
    • How to write justifications that are specific, grounded, and consistent with your rating
    • Each of the dimensions you'll judge: Instruction Following, Visual Quality, AI Appearance, Correctness, and the two Preference ratings
    • Best practices for staying accurate and consistent across a task

    The final sections put everything together with hands-on practice: spotting image oddities, evaluating how well an image follows a prompt, and comparing two images end to end. Passing the last knowledge check unlocks your tasking.

You’ll receive an email to your personal inbox with your Slack invitation. If you don’t see it, check your spam folder. Log in via Slack.com or the app.

You’ll be assigned to a pod channel, which is your primary space for day-to-day communication and support.

Once you receive your Slack invitation at your Workada email, log in via Slack.com or the app. You will be assigned to a pod channel: this is your primary space for day-to-day communication and support.

Pod Lead assignment happens automatically after you activate Slack. If you aren’t assigned to a pod within the next business day, email support@op.workada.co.

It's used to sign into Slack: there's no standalone inbox. All mail is automatically forwarded to your personal email, so please check your personal inbox regularly (including your spam folder) for any communications from Workada or your Pod Lead.

This is a known issue: some antivirus configurations interfere with the timer.

Use Google Chrome. Don’t use Safari or Firefox, they don’t reliably track your timer activity and can log active work as idle.

iOS, Android, and tablets, Chrome OS and Chromebooks, and Linux aren’t supported. An installed antivirus or the Firefox browser can also interfere with the timer. On a MacBook, use only Google Chrome.

Check the Project Resources tab, which covers most project-specific answers. If you can’t find what you need there, ask your Pod Lead.

B

Payment

As pay may vary by project, your rate is reflected on your dashboard under the Payments tab. You can find it at any time in the Payments tab on your dashboard.

Your payout is calculated as: (Eligible Hours × Hourly Rate) + Incentives (if earned)

"Eligible Hours" is the time that counts toward payment under our Pay Policy: this excludes idle time, activity on non-approved sites, and training time beyond your project's cap. See the Pay Policy and Timer tab for the full breakdown of what's eligible.

You can verify your expected payout at any time by checking your session history in the Sessions tab on your dashboard before raising a support ticket.

Payments are processed three times a week. The day your work is processed depends on when you complete it:

Finish your work byProcesses on
Sunday, 11:59 PM PTMonday
Tuesday, 11:59 PM PTWednesday
Thursday, 11:59 PM PTFriday

Once processed, Stripe typically takes 1–3 business days to deposit the funds into your account. If a processing day falls on a public holiday, payment will be moved to the last business day before it.

Yes. Workada never sees or stores your payment information. Your banking details are held directly and securely by Stripe, one of the world's largest payment processors.

Check your session history in the Sessions tab first to verify your logged hours. If something still looks off, email support@op.workada.co with the details.

Contributors are classified as independent contractors and are responsible for managing their own tax obligations. Workada does not withhold taxes on your behalf. Tax documentation is generated automatically through Stripe.

After you refer someone through your dashboard, your referral incentive shows up in your Payments tab within about two pay cycles. The clock starts once the person you referred completes their off-platform task.

C

Time tracking

All work must be tracked using the Workada Timer. You are responsible for clocking in before you start working and clocking out when you stop.

The timer also manages itself in a few ways: it automatically pauses after about five minutes without mouse or keyboard input, and it pauses whenever you're on a site or app that isn't approved for your project: you won't be able to resume until you switch back. Supporting tools like Slack and Google Search are fully paid as long as they're on the approved list, and training/calibration time is paid up to a cap that varies by project. For the full breakdown of what counts as paid time, see the Pay Policy and Timer tab.

Workada Timer app showing Start and Stop buttons
Workada Timer app
Workada overlay mini-timer
Overlay mini-timer

You can see all of your tracked time in the Activity tab of your Workada dashboard, broken down by tasking versus other time, by day, and by the sites and apps you used.

The Activity tab of the Workada dashboard: Tasking and Other time totals, a per-day time-tracked chart, and a breakdown of most-used websites and apps.
The Activity tab in your dashboard, where you can review your tracked time.

Training and calibration are paid up to a cap specific to the project you have applied to. Once you pass that cap, the timer shows a warning and any time clocked in after that isn’t paid. If you run into this, reach out to your Pod Lead.

Workada timer showing a Training Time Exceeded warning: 27:56:46 logged against the 04:00:00 paid training cap, noting additional training time is no longer eligible to be paid.
The timer’s Training Time Exceeded warning once you pass the time cap.
✓ Do
  • Click Start before you begin any work or training
  • Stop the timer when you finish your session
  • Add a manual session if you forgot to start the timer
  • Request a session edit if the timer recorded the wrong time
✗ Don't
  • Start working without the timer running: time not tracked cannot be paid
  • Close the timer window thinking it stopped: it continues running in the background until you click Stop
  • Run the timer when you're not actively working
  • Forget to start the timer for training: it's paid time too
  • Submit duplicate manual sessions for the same time period

Start your timer any time you are doing work-related activities: this includes training, tasking, attending meetings with your Pod Lead, or anything else related to your work on the platform.

Supporting tools like Slack, Google Search, and meetings are fully paid as long as they're on the approved tools list: see the Pay Policy and Timer tab for the full list. Note that AI assistants are not an approved tool for producing task output; using one won't count toward paid time and may affect your standing on the project.

Go to Dashboard > Sessions > + Add Manual Session. Enter the time window you worked and a brief description of the work completed. Submit it for review.

Time added manually is still paid if approved by Workada.

Manual sessions that are approved by Workada are paid out in the next scheduled payment cycle.

Yes: you can connect the Workada Timer on multiple computers if you work from more than one. To download it, go to the Sessions page or the Settings page on your dashboard and look for the download link. Note that you cannot run the timer on two computers at the same time.

Go to Dashboard > Sessions, find the incorrect session, and click "Request Edit." Adjust the time to what's correct, enter the reason, and submit. You can also contact your Pod Lead or email support@op.workada.co with the session details.

Check your Sessions tab at workada.com/sessions: your time is often saved on the backend even if it disappears from the app view. If the session is missing, add it manually via + Add Manual Session and contact support@op.workada.co with the details.

Go to Dashboard > Sessions, find the affected session, and click "Request Edit." Adjust the time to what's correct, enter the reason, and submit. If the session is missing entirely, add it manually via + Add Manual Session. If the issue persists, email support@op.workada.co with the details.

Neither. There's no minimum or maximum: you work at your own pace. What we care about most is the quality of the work you submit, not the number of hours you put in.

That said, when we review metrics, we do look at whether the time spent on a task reflects the quality of the work. Rushing through tasks is something we're able to identify, so take the time you need to do the work well.

We do look for Contributors who keep tasking regularly. Inactivity means going quiet on actual work, no submitted tasks for a stretch, not a quiet week in Slack or a slow reply to a message. If you need to step away from tasking for 7 or more days, let your Pod Lead know in advance.

Take the time you need: just let your Pod Lead know ahead of time so they can plan around your absence.

Use Stop when you're done working for the session; it ends your session and records the time. Use Pause for a short break you'll come back from; it keeps your session open and removes the break from your paid hours, so you don't have to stop and restart.

Both are fine to use. Pause just saves you from juggling separate start and stop periods for short breaks. Note that your balance won't show up in your Sessions history until you click Stop, and that paused time isn’t eligible for paid time.

This is a known issue: a few Contributors have had the chime keep ringing even after selecting "Still working." We're looking into a fix. In the meantime, opening the main timer window stops it.

We're aware that paused breaks can appear labeled as manual adjustments on your activity chart. We're working to ensure they're correctly labeled as ineligible time instead. This is a display issue only; your paused time is still correctly excluded from paid hours.

If you're on a MacBook, please use Google Chrome for tasking. We've confirmed that Safari doesn't reliably track activity on Mac; it can show active work as "Other" or "no activity detected," even when you're actively tasking. Switching to Chrome will track your time correctly and consistently.

D

Offboarding

Offboarding means losing access to continue contributing at Workada. There are five reasons you may be removed from the platform:

  1. Low quality work: if your quality falls below an acceptable threshold. Coaching always comes first, so reach out to your Pod Lead to run shadow sessions and improve your quality before it gets to this point
  2. Dishonesty: using AI to do your work for you, copying/pasting someone else's or your own prior work without doing the task yourself, spam, or time milking. First incident results in a warning; second results in offboarding
  3. Inactivity: going 7+ days without tasking (no submitted work) and without giving your Pod Lead a heads-up. This is measured by your tasking, not by your Slack activity or whether you reply to every message, and planned time off you've flagged in advance doesn't count.
  4. Community guidelines violation: egregious violations result in immediate offboarding with no warning
  5. Failing training or calibration: failure to pass either within the required attempts may result in offboarding

In all cases, you will be paid for all work completed up to your offboarding date, processed in the next scheduled payment cycle. Your Dashboard will remain available so you can check your Sessions and Payments history.

Notify your Pod Lead directly that you'd like to leave the platform. They will flag your departure to Workada Operations, who will revoke your Slack and Workada email access. Your Dashboard will remain available so you can check your Sessions and Payments history. You will be paid for all work completed up to your offboarding date in the next scheduled payment cycle.

✉️

Still have questions?

Email support@op.workada.co with your full name, a clear description of your issue, and any relevant details (task URLs, session times, screenshots, etc.) so the team can assist you efficiently.

At Workada, everything we build, every project, every partnership, every platform decision, comes back to three things. These aren't values we put on a wall. They're the principles that shape how we work, what we reward, and where we're headed together.

Quality

Quality is our north star. We don't trade great work for speed or scale. Every annotation matters because it shapes real AI outcomes, and because it determines what Contributors can access, earn, and achieve on the platform.

Quality isn't a standard we impose. It's the key that opens every door. The more you bring it, the more Workada brings back: more projects, more opportunity, more room to grow.

Community

Workada isn't just a platform, it's a community. We look out for our Contributors, and they raise the quality of everything we build. That's not a tagline; it's how this actually works.

When Contributors grow, in skill, in earnings, in impact, Workada grows too. We invest in the people who show up, do the work, and push each other to be better. This is a community that improves itself, together, and there's always room for people who want to be part of that.

Opportunity

At Workada, there's no ceiling on what you can achieve. The opportunity to grow, as an individual, in your career, and within the company, is real and it's ongoing.

High-quality Contributors always have projects waiting. Downtime shouldn't be something you worry about: that's our job. The pipeline stays full for those who keep the bar high. Show up, deliver, and the doors keep opening. That's the Workada promise.

Workada Time Tracking & Pay Policy Update

This policy explains how working time is tracked and paid. The principle is straightforward: you're paid for time spent actively doing approved tasking work. Below is how time is categorized and what counts toward payment.

Timer walkthrough

Here’s a short walkthrough of the timer and how your paid time is tracked.

What Counts as Paid

Project Tasking

Payment covers project tasks performed within the agreed and approved scope of your project. This work should take place on the approved tasking sites and apps.

If you are unsure whether an activity or tool is approved, consult your project instructions or contact your Pod Lead.

Training & Calibration

Onboarding time is paid up to a cap, specific to the project you have applied to. Workada's timer will show warning states as you approach the limit. Time spent on training and calibration beyond the cap is unpaid.

Workada timer showing a Training Time Exceeded warning: 27:56:46 logged against the 04:00:00 paid training cap, noting additional training time is no longer eligible to be paid.
The timer shown above is from a project with a 4-hour cap for training and calibration. When you pass the cap, the timer shows this warning, and any time logged past it isn’t paid.

Supporting Tools

Certain tools support the work without constituting the work itself. Time spent on these tools is paid. The approved supporting tools are:

  • Slack
  • Zoom
  • Google Meet
  • Google Search
  • Word
  • Google Docs
  • Email
  • Image hosting sites used by tasks (for opening a task's images in a new tab to zoom or compare)
Your timer automatically pauses when you visit a denied site, for example amazon.com, tiktok.com, chatgpt.com, or linkedin.com. It picks back up when you return to approved work.

Use of AI Tools

Task output must be your own work. You may not use AI assistants (for example ChatGPT, Claude, Gemini, Copilot, or similar) to generate, complete, or fill in any part of your task deliverables. AI use will not be counted toward paid time.

If a specific project permits AI use for a defined purpose, that will be stated explicitly in your project instructions. Absent that, assume AI tooling is not allowed for producing work. Violations are treated as a quality and integrity issue, not a pay issue, and may affect your standing on the project.

Idle Time

To ensure that paid time reflects active work, the timer pauses automatically after five minutes without mouse or keyboard input, on any screen, application, or tab. When the timer pauses, you will be asked to confirm that you are still working. Acknowledging the prompt resumes tracking. Paused time is not paid.

Ineligible Activity

All other activity falls outside paid work, such as shopping, food ordering, streaming, and general personal browsing. The timer pauses automatically whenever you are not on an approved site, and that time is not paid. You will not be allowed to resume the timer until you switch to an approved tasking site or app.

Seeing Your Eligible Time

Your Workada timer will now separate tracked time into two categories, so you can see as you work what will and will not count toward payment:

  • Eligible time: approved core tasking, supporting-tool time, and training time within your cap. This is the time that counts toward payment, subject to the standard review applied to all submitted work.
  • Not eligible: idle time, activity on non-approved sites, and training time beyond your cap.

This breakdown is visible in your timer at any point during a session, so there are no surprises at the end of a pay period. If a block of time is marked not eligible and you believe that is an error, contact your Pod Lead.

Supported Browsers

We recommend using Google Chrome when tasking on approved project sites. If you're on a MacBook, Chrome is required: we've confirmed that Safari doesn't reliably track activity on Mac, and can show active work as "Other" or "no activity detected" even when you're actively tasking. Use of other browsers may affect the accuracy of your timer session.

Where you can start the timer

Your timer can only start from an approved work location. If you’re in a restricted area, the timer won’t start and you’ll see the message below. Restricted locations are set per project, so what counts as approved depends on the project you’re on.

VPN use isn’t allowed. A VPN can make your location look different from where you actually are, which can block your timer or flag your account.

Workada Timer showing a Can't start the timer here error: timers can only be started from an approved work location.
If you’re outside an approved location, the timer won’t start.

Timer Best Practices

The Workada Timer is how your work turns into pay, so it's worth using well. Here's what counts as active work, what doesn't, and what to do if something looks off.

The Timer, and What Counts

Every Contributor must download the Workada Timer, which you'll find on the Workada Dashboard. Run it whenever you're actively working, and active work covers more than tasking:

Tasking Training & calibration Tutorials Work-related Slack Googling & researching for a task Work meetings & huddles Reading & thinking through a task

Stop vs. Pause

Use Stop when you're done working for the session; it ends your session and records the time. Use Pause for a short break you'll come back from; it keeps your session open and removes the break from your paid hours, so you don't have to stop and restart.

Both are fine to use. Pause just saves you from juggling separate start and stop periods for short breaks. Note that your balance won't show up in your Sessions history until you click Stop, and that paused time isn’t eligible for paid time.

What the Timer Tracks

  • Records which application or website is in focus, not what's on your screen
  • Doesn't take screenshots, read window titles, or capture what you type
  • Tracks mouse/keyboard input timing only, never content
  • Records the country of your IP at session start
The green light on the timer means you’re on an approved site. It doesn’t indicate productivity, it simply confirms the site you’re visiting is allowed.

Task-Level Detail

  • On Workada and your project’s webpage, the full page address is recorded and linked to the task you opened
  • On every other site, only the site name is recorded (e.g. "wikipedia.org"), never the specific page

Overlay Colors

  • Green means the focused window counts as eligible time
  • Gray means it doesn't
  • Eligible time includes approved tasking sites, supporting tools (Slack, Zoom, Google Meet, Google Search, Word, Google Docs, Email), and training time up to your cap

Switching Windows Briefly

  • Quick switches won't cost you time thanks to a short grace period before auto-pause
  • Extended time on non-approved sites pauses the timer and isn't paid
  • Think a block was marked wrong? Contact your Pod Lead

Automatic Pausing

  • Pauses on its own after five minutes with no mouse or keyboard input
  • You'll be asked to confirm you're still working; acknowledging it resumes tracking

Do's and Don'ts

Do

  • Run the timer only while actively working: anything in the list above counts.
  • Pause the timer the moment you step away for anything personal, even a few minutes.

Don't

  • ×Leave the timer running during personal browsing, socializing, or idle waiting.
  • ×Treat Slack as unlimited paid time. Only work-related Slack belongs on the timer: chat and socialize on Slack once the timer is off.

Additional Guidelines

  • You can check your hours yourself on the Dashboard's Sessions tab. Pay is hours worked × your rate.
  • Forgot to start your timer while working? You can add a manual session in your timesheet.
  • Left the timer running by accident while not working? You can edit your session to remove the time.
  • If a warning email lands, there's no need to panic or stop working. Make sure you're using the timer when you're supposed to: document your side, and escalate through your Pod Lead rather than stress.
  • One correction request is enough: a single support email or Sessions-tab request does the job, and duplicates slow things down.

Known Issues

  • Chime not stopping after you confirm: a few Contributors have had the idle-check chime keep ringing even after selecting "Still working." We're looking into it; opening the main timer window stops it in the meantime.
  • Paused time showing as a "manual adjustment": paused breaks may currently be labeled as manual adjustments on your activity chart. We're working to ensure they're correctly labeled as ineligible time instead. This is a display issue only; your paused time is still excluded from paid hours.

0 of 3 reviewed

We've confirmed that Safari doesn't reliably track activity on Mac; it can show active work as "Other" or "no activity detected," even when you're actively tasking.

If you're on a MacBook, please switch to Google Chrome for your timer to track correctly and consistently.

Referrals just got easier! You can now submit and track referrals directly from your Workada Dashboard: no separate form needed.

Here’s how it works:

  • Once you’ve completed 5 hours of work on the platform, the Referrals section will appear automatically in your Dashboard.
  • Submit your referrals there and follow their progress in real time. All tracking now lives in the platform, so you’ll always know where things stand.
  • Your referral gets an automatic email inviting them to sign up. Once they’re approved and complete their first project task, you both earn a reward.

Rewards at a glance:

  • Refer a Contributor: you earn $50 and they earn $20.
  • You can have 5 referrals at a time, and the limit refreshes every 30 days.

One more thing: we’re retiring the old referral form, so please use your Dashboard for all new referrals going forward. Any referrals you’ve already submitted through the form will still be tracked: nothing gets lost.

Thank you for helping grow this community. Happy referring!

You’ll find announcements from the Workada team on this tab: updates, news, and changes involving the Workada Contributor community. Try to make it a habit to check here before you start tasking. Once you’ve read an announcement, mark it as reviewed: new ones will arrive unmarked.

Everything for the Multimango and Video Eval project: your Slack channels, how to evaluate image tasks, the video training and calibration, and project-specific questions.

Channels

0 / 0 checked
1 · Understand the prompt
  • Read the prompt carefully: including implied context like "for a school project" or "should look like an ad"?
  • Listed what must be included: objects, people, text, layout, colors, style?
  • Noted what shouldn't have changed from the prompt image: did anything change anyway?
2 · Compare A vs. B: and the prompt image
  • Zoomed into both images (hands, faces, text, edges, backgrounds), not just glanced at the whole?
  • You're comparing A vs. B, not just describing one of them?
  • Checked whether either output changed something the prompt never asked to change?
3 · Ask the four questions
  • Does it follow the prompt? Everything requested is there, in the requested style, layout, and text: and the rest is untouched?
  • Does it make sense? Logical scene, realistic proportions, natural behavior, sensible context?
  • Does it look polished? Could you recreate this on purpose in Canva or Photoshop: or is it distorted, messy, obviously AI?
  • Is it correct? Objects, anatomy, logos, text all accurate (verified against reality where checkable), judging only what applies to this prompt?
4 · Hunt for AI artifacts
  • Anatomy: distorted or misshapen bodies, missing or extra limbs, merged objects, impossible proportions?
  • Text: garbled, misspelled, or gibberish lettering anywhere in the frame?
  • Repetition: cloned faces or duplicated objects hiding in crowds and backgrounds?
  • Physics: unrealistic textures, lighting, reflections, or shadows?
5 · Write a strong justification
  • You compared both images and explained why one wins, with specific examples: no vague statements?
  • You used the prompt's own wording and descriptive language: distorted, misshapen, cohesive, realistic, whereas?
6 · Final consistency check
  • Your comments support your ratings: one consistent case from start to finish?
  • Your Overall Preference matches the issues you described in your comments?
  • Proofread, inside the character limit, timer running, session confirmed?
01
Specific
Name the object, the error, the location: not just "A looks better."
02
Right category
AI artifacts under AI Appearance first; blur and composition under Visual Quality first.
03
Consistent
Your Overall Preference and your body text point to the same winner.
04
Prompt-built
Every observation ties back to what the prompt actually asked for.
PromptUpdate this image to a winter Christmas holiday scene. Change the four individuals outside the picket fence to carollers. Change the person in the doorway to wear a holiday sweater with a Santa hat.
The two outputs, A and B, side by side
The two outputs: A vs. B
A wins · 5/5Audit: pass
Overall → A
A builds a cohesive winter Christmas scene. Both convert the four people to carollers and add the holiday sweater: but B keeps the pumpkins from the original, which aren't Christmas themed.
Instruction Following → A
Both satisfy every explicit request. Though not explicitly stated, A removes the pumpkins from the source; B does not: the intent of "winter Christmas scene" covers it.
Visual Quality → A
A's composition is more balanced without the pumpkin clutter, and its carollers are spaced naturally where B's are bunched up. Both use soft natural night light.
AI Appearance → A
A reads naturally with Christmas lights, snow, and winter clothing. B's pumpkin decorations look misplaced against the snowy backdrop. Both preserve the source composition.
Correctness → A
A is a more correct holiday scene; B is less correct for keeping the pumpkins. Both correctly show carollers holding song books and the doorway figure in Christmas clothing.
Key lessonWhen both images do something right, say so: balanced credit makes your verdict more credible. And when you spot something the prompt didn't explicitly state, reason through whether the prompt's overall intent covers it.
PromptRemove the food from the table and replace with brownie with ice cream dessert and coffee. Preserve the people and scene context.
The two outputs, A and B, side by side
The two outputs: A vs. B
A wins · 5/5Audit: pass
Overall → A
Neither image is perfect, but A better follows instructions, is more correct, and looks less AI-generated than B.
Instruction Following → A
Both remove the food, but A preserves the people and scene better: in A each person has at least one dessert and all but one have coffee. In B, half the people are missing both.
Visual Quality → A
Zoomed in, A has better lighting consistency, resolution, and sharpness, and a cleaner, more pleasing composition.
AI Appearance → A
In B, half the people are missing dessert and coffee, the girl with glasses is missing a hand, the fingers holding the fork are unnatural, and the ketchup-bottle text is garbled.
Correctness → A
Neither image has the correct number of desserts and coffees, but A is closer: and it more correctly preserves the people and scene context.
Key lessonCounts are evidence. "Half the people are missing dessert and coffee" beats "some items are missing": and honest hedging ("neither image is perfect," "slightly better") reads as careful judgement, not weakness.
PromptTurn this photo into an editorial personal headshot. Maintain the integrity of the lady's face, expression, pose, and identity. Change the background to grey and use a clamshell lighting effect.
The two outputs, A and B, side by side
The two outputs: A vs. B
B wins · 5/5Audit: pass
Overall → B
B has better clamshell lighting; A leaves too many shadows on the lady's face. Both maintain her face, expression, pose, and identity, and both have a grey background.
Instruction Following → B
B gets the requested clamshell lighting right. Both otherwise read as editorial headshots and keep the grey background and her identity intact.
Visual Quality → B
B's lighting is more even and it's slightly sharper. Composition is similar in both.
AI Appearance → B
In A, the model changes her hair, adds a collar, and swaps the buttons on her jacket. Otherwise both look natural without significant artifacts.
Correctness → B
B's clamshell lighting correctly eliminates facial shadows. A shadows her face, changes her hair, collar, and jacket buttons, and darkens her sunglasses.
Key lessonSpecificity goes all the way down: not just naming what changed, but where exactly, and why it matters for the prompt. When another reviewer might see it differently, acknowledge it and explain your reasoning. That's what separates a confident 5 from a vague one.
The goal isn’t simply to find mistakes: it’s to decide which image does a better job of fulfilling the user’s request.
Rating guideline: don’t use AI tools to evaluate tasks or check facts. Google is fine for finding reliable sources when fact-checking is needed.
A much betterA slightlyTieB slightlyB much better
Quick reference
QuestionAsk yourselfLook for
Overall PreferenceWhich image is better overall?Prompt fit, quality, and errors weighed together
Instruction FollowingDid it do what the prompt asked?Missing items, wrong style, wrong colors, extra clutter
Visual QualityIs it well presented?Blur, lighting, cropping, focus, composition
AI-Generated AppearanceDoes it have AI-looking mistakes?Extra fingers, warped faces, melted objects, gibberish text
CorrectnessAre checkable details right?Spelling, dates, labels, counts, maps, facts
The five questions, in depth
Overall Preference
Which response do you prefer overall: and why?
What to look for
  • Which image best fulfills the user’s request overall
  • Which strengths and weaknesses matter the most
  • Whether one major issue outweighs several minor ones
  • Overall visual appeal, clarity, and presentation
Keep in mind
  • This is a holistic judgment made after your detailed review, weighing everything you found across all dimensions: not a quick first impression, unless the task instructs otherwise
  • Not every category carries the same weight in every task
  • It’s not a category count: give a well-supported answer rather than tallying which image “won” more questions
  • You’ll rarely have a true tie
  • Avoid ties unless the images are genuinely very similar overall
ExampleImage B is better overall. Although A has slightly sharper details, B follows the prompt more closely and doesn’t contain the major AI-generated artifacts found in A.
Instruction Following
Which image better follows the prompt?
What to look for
  • Does it match the requested style: realistic photo, illustration, edit, greeting card?
  • Are the requested objects, people, attributes, and edits all there?
  • Are any required details missing or incorrect?
  • Were unnecessary or conflicting elements added?
  • For editing tasks: were unchanged parts of the original preserved?
Keep in mind
  • Start with the big picture: what was the user actually asking for?
  • Not every requirement carries the same weight; think about which issues have the biggest impact
  • Added details aren’t automatically wrong if they fit naturally and don’t conflict with the prompt
  • Judge the user’s overall intent, not just a checklist
ExampleResponse A follows the instructions much better. It includes the key details, such as the address and RSVP number, while Response B leaves out some of that information and is missing the balloon that says “Happy Birthday.”
Visual Quality
Which response has better visual quality?
What to look for
  • Sharpness and focus
  • Lighting and contrast
  • Composition and framing
  • Resolution and texture quality
  • Realistic perspective and depth of field
  • Editing seams or blending issues, when applicable
Keep in mind
  • Look at the whole image first, then zoom in: small issues can matter
  • Blur isn’t automatically a flaw; motion blur or shallow depth of field may be intentional
  • A technically sharp image isn’t always better if composition or presentation is weaker
  • Don’t over-penalize small flaws that don’t significantly affect the image overall
ExampleResponse B has better visual quality. A has blurry details (some of the cups on the table), and the chair is cut off by the frame.
AI-Generated Appearance
Which response looks less AI-generated overall?
What to look for
  • Distorted anatomy, especially hands and fingers
  • Warped or asymmetrical faces
  • Gibberish or malformed text
  • Duplicated elements and repetitive patterns
  • Melted or merged shapes; objects that didn’t form properly
  • Impossible geometry or perspective; unnatural textures
Keep in mind
  • Focus on artifacts, not on which image looks more realistic
  • One major artifact may matter more than several smaller flaws
  • Stylization alone doesn’t mean an image is AI-generated: illustrations and cartoons can be fully coherent
  • If neither image has noticeable artifacts, don’t hunt for issues that aren’t there; an equal rating may be appropriate
ExampleResponse A appears less AI-generated. Although it isn’t as sharp, Response B contains gibberish text and extra fingers, which are more serious issues.
Correctness
Are there details that can be checked: and are they right?
What to look for
  • Incorrect dates, numbers, or measurements
  • Wrong labels, signs, or logos
  • Spelling errors
  • Impossible layouts or spatial relationships
  • Historical, geographical, or domain-specific inaccuracies
Keep in mind
  • Use Google to verify when needed: correctness is about details that can be objectively checked
  • Missing requested elements belong under Instruction Following
  • Gibberish or unreadable text usually belongs under AI-Generated Appearance
  • If there are few checkable details, this question may not be a major factor for that task
ExampleResponse A is much better because the Adidas logo in Response B is incorrect: believable at first glance, but it doesn’t match the real version.
Where does this observation belong?

Some observations could reasonably fit more than one question. Think about the main issue you’re describing and place it where it fits best.

If the issue is…It usually belongs under…
Prompt asks for 5 people, image shows 4Instruction Following
Prompt asks for red, image shows blueInstruction Following
Sign says “July 32, 2026”Correctness
“Miami” is spelled “Maimi”Correctness
Text is unreadable or random symbolsAI-Generated Appearance
A hand has seven fingersAI-Generated Appearance
Image is blurry or pixelatedVisual Quality
Subject is cropped too tightlyVisual Quality
A person is missing from the scene entirelyInstruction Following
A logo looks almost right but isn’tCorrectness
Fabric and skin have blended togetherAI-Generated Appearance
The same face appears 3 times in a crowdAI-Generated Appearance
Shadows point in conflicting directionsAI-Generated Appearance / Visual Quality
A building is architecturally impossibleAI-Generated Appearance / Correctness
Same date, three different homesUnreadable or warped date → AI-Generated Appearance. Readable but impossible (“February 32”) → Correctness. Readable and valid but not what the prompt asked for → Instruction Following.
When both applyYou can mention the same observation in two categories if each justification covers a different part. “Maaimi” in warped lettering → Correctness: the city name is misspelled. AI-Generated Appearance: the text is distorted and poorly rendered.
Stay attentive to the details
  • Keep the prompt in mind for every question. User intent is your context: did they ask for an illustration, a realistic image, a card, a studio shot? Bright sunlight or a rainy day? Warmth, sadness, shock?
  • Zoom in: every time. You can’t evaluate these tasks properly at full-frame view. Small issues decide close calls.
  • Google what you don’t know. If the prompt mentions something unfamiliar, look it up on official sites. Answering with verified background knowledge makes your data far more accurate and valuable.
  • 250 characters go fast. You won’t fit every detail, so lead with the ones that best support your answer for that question.
  • Keep your answers separate. The questions are mostly about different things, so different details should support each one: avoid repeating the same observation everywhere.
  • Avoid circular reasoning. “It looks more realistic because it’s more realistic” isn’t a justification. Name the tells: the plasticky plants, the gibberish text on the book.
Where the tells hide when you zoom
Light & shadow
Crushed blacks, blown-out highlights, impossible light directions, missing shadows, reflections that don’t match the scene.
Edges & seams
Visible editing seams, inserted objects that look superimposed, mismatched grain, focus that changes where it shouldn’t.
Surfaces
Waxy or over-smoothed skin, flattened textures, matte rendered as glossy, tiled or stamped repeating patterns.
Text
Misspelled, truncated, or repeated words, garbled scripts, wrong language or reading direction, unpreserved fonts.
Color & exposure
Over- or under-saturation, banding, hue shifts on specific objects, inconsistent color temperature across regions.
Logic & scale
Implausible object sizes, floating objects, clipping, perspective failures: things that couldn’t exist in the scene.
When can it be a tie?
Rarely a tieOverall Preference, Instruction Following, and Visual Quality almost always have a winner. If they feel tied, look closer: zoom in, re-read the prompt, and compare the details again.
A tie can be rightAI-Generated Appearance can tie when neither image has noticeable artifacts. Correctness can tie when there’s nothing to check: or both images get the checkable details right. If it’s truly a tie, say so and explain why.
Instruction Following vs. Correctness

Instruction Following asks: did the model do what the prompt asked? Correctness asks: is what the model created factually accurate? An image can pass one and fail the other: here’s how the same scenario scores on both.

ScenarioInstruction FollowingCorrectness
Prompt asks for a portrait of a man in a blue coat; the image shows a black coat. Error
The prompt specified blue and the model ignored it.
OK
The image itself is accurate: the model just didn’t follow the requested color.
Prompt asks for a map of the Caribbean with its islands labeled; the map appears, but some labels are wrong. OK
A labeled map of the Caribbean is what was asked for, and that’s what was generated.
Error
Some island labels aren’t factually accurate.
Prompt asks for a square Save-the-Date for Juan & Mitchell reading “Save the Date – October 13th”; the layout is right but it says “Michelle” and “Octobre.” Error
The prompt asked for specific names and dates, and the model failed that request.
Error
The name and month are misspelled.
Prompt asks for a chess set with red and white pieces; the image shows a correct chess set with black and white pieces. Error
The prompt specified red and white pieces and the model ignored the requirement.
OK
The chess set itself is correct.
Prompt asks for a bird’s-eye view of Times Square at night; the setting and view are right, but buildings are misshapen and signage has text errors. OK
The model created the correct setting and time of day.
Error
The malformed buildings and text errors need to be pointed out.
These are simplified examples: real tasks can have multiple errors in either category, or both. Read the prompt thoroughly to check every requirement for Instruction Following first, then analyze the images for Correctness errors.
Don’t hunt for a single flaw: look for an accumulation of small improbabilities. Real photos have messy coherence; AI images tend to have polished inconsistency.

A lot of AI images look convincing at first glance, but they almost always contain subtle mistakes: and your job is to find them. Don’t just ask whether an image looks good. Ask what’s wrong with it, and explain it clearly.

1 · Unnatural anatomy
Extra, missing, or fused fingers and toes; hands that are too smooth, too large, or bent impossibly; missing limbs; asymmetrical, duplicated, or melted facial features; fused teeth; ears merging into hair; skin with no pores or imperfections.
Zoom in on every hand and face: these are where AI struggles most, and they’re easy to miss at full view.
2 · Distorted or nonsensical text
Gibberish, random symbols, or near-words (“COFFFE,” “Rstrnt”); letters that melt or bleed together; text that falls apart when you zoom in; almost-correct logos; garbled tattoos or clothing text; numbers that don’t add up.
Gibberish belongs under AI-Generated Appearance. Readable but factually wrong (like a misspelled real name) is Correctness.
3 · Unnatural textures
Airbrushed, waxy, or plastic skin; fur or hair with no individual strands; wood, stone, or brick that repeats or looks printed on; fabric folds that don’t behave like fabric; metal, glass, or water with inconsistent reflections.
Ask: does this fabric move like fabric? Does this wood grain make physical sense?
4 · Lighting & shadow inconsistencies
Shadows pointing in different directions; reflections that don’t match the scene; objects casting no shadow when everything else does; skin glowing unnaturally evenly; highlights on the wrong side of a face.
Lighting issues often sit under Visual Quality: but if they make the scene physically impossible, they belong under AI-Generated Appearance.
5 · Merged or melted objects
Watch straps merging into wrists; clothing blending into skin; cups merging into tables; furniture legs disappearing into floors; hair merging into the background; glasses arms vanishing into hair.
Check every edge where two objects meet: it’s easy to miss at full view.
6 · Impossible geometry & perspective
Staircases that lead nowhere; rooms with impossible layouts; objects at the wrong scale relative to each other; structurally impossible windows and doors; vehicles missing wheels or doors; distorted background figures.
Squint at the full image first. If something feels spatially off, trust that instinct: then zoom in to confirm.
7 · Unnatural repetition
Several people in a group with near-identical faces; the same figure appearing twice in a crowd; tiles or wallpaper repeating in an obvious pattern; visibly duplicated clumps of leaves or flowers; identical animal markings.
Pay close attention when a prompt asks for multiples: four children, a crowd, a row of objects. Each one should be distinct.
8 · The “too perfect” AI look
Extreme depth of field on ordinary scenes; dramatic rim lighting on everyday subjects; movie-poster color grading; exaggerated “stock” emotions; perfect symmetry real photos never achieve; zero messiness anywhere in frame.
Matters most for photorealism tasks: if the prompt asked for natural or candid, over-polish is worth noting under AI-Generated Appearance.
Use specific photographic terminology whenever possible. Avoid “looks weird” or “bad quality”: name the visual characteristic that’s actually driving your score.
meets the bar situational / intentional × flag as an issue
Sharpness & focus
  • Sharp: the subject is clear and well-defined. “The image is sharp, and facial details are clearly visible.”
  • Crisp: fine details render cleanly with strong clarity. “The feathers appear crisp with excellent texture detail.”
  • Soft: details are slightly smooth or less defined, not necessarily wrong. “The image is slightly soft around the edges of the subject.”
  • ×Blurry: important details are hard to see due to lack of focus. “The image appears blurry, particularly around the face.”
  • ×Out of focus: the focus is placed incorrectly. “The subject is out of focus while the background appears sharper.”
Exposure
  • Properly exposed: highlights and shadows are balanced. “The image is properly exposed and retains detail in both bright and dark areas.”
  • ×Overexposed: too bright, losing detail. “The sky is overexposed and lacks visible cloud detail.”
  • ×Underexposed: too dark to distinguish details. “The subject’s clothing is underexposed and difficult to see.”
  • ×Blown highlights: bright areas lose all texture. “The highlights on the white shirt are blown out.”
  • ×Crushed shadows: dark areas lose visible detail. “The shadows beneath the table are crushed and contain no visible detail.”
Lighting
  • Natural lighting: light appears realistic and believable. “The natural lighting creates a realistic outdoor appearance.”
  • Soft lighting: smooth transitions, minimal harsh shadows. “Soft lighting creates flattering skin tones.”
  • Well-lit: the subject is clearly visible. “The image is well-lit and easy to evaluate.”
  • ×Harsh lighting: strong shadows and intense highlights. “Harsh lighting creates distracting shadows on the face.”
  • ×Flat lighting: the image lacks contrast and depth. “Flat lighting causes the subject to blend into the background.”
  • ×Poorly lit: insufficient or inconsistent lighting. “The scene is poorly lit, making details difficult to assess.”
Contrast & tonal range
  • Strong contrast: clear distinction between light and dark. “The strong contrast helps the subject stand out.”
  • Balanced contrast: light and dark areas are well controlled. “The image demonstrates balanced contrast throughout.”
  • ×Low contrast: tones lack separation. “The image appears dull due to low contrast.”
  • ×Muddy tones: tones blend together without definition. “The background contains muddy tones with little separation.”
Color
  • Accurate color balance: colors appear realistic and natural. “Skin tones exhibit accurate color balance.”
  • Vibrant: colors appear rich and lively. “The vibrant colors enhance visual appeal.”
  • Muted: colors appear intentionally subdued. “The muted palette creates a softer mood.”
  • ×Oversaturated: colors are excessively intense. “The flowers appear oversaturated and unrealistic.”
  • ×Undersaturated: colors are washed out. “The image looks undersaturated and lacks visual impact.”
  • ×Color cast: an unwanted tint affects the image. “The image has a noticeable blue color cast.”
Texture & detail
  • Rich texture: surface details are clearly visible. “The wood grain exhibits rich texture.”
  • Fine detail: small details survive a close look. “Fine detail is preserved in the hair.”
  • Clean rendering: objects appear smooth and defect-free. “The image demonstrates clean rendering throughout.”
  • ×Artificial smoothing: details are unnaturally softened. “Artificial smoothing reduces skin texture.”
Depth & dimension
  • Good depth: a realistic sense of space. “The image demonstrates good depth between foreground and background.”
  • Strong subject separation: the subject stands out clearly. “Strong subject separation draws attention to the focal point.”
  • Shallow depth of field: sharp subject, blurred background. “The shallow depth of field effectively isolates the subject.”
  • ×Flat appearance: the image lacks dimensionality. “The scene appears flat due to limited depth cues.”
Composition
  • Well-composed: elements are arranged effectively. “The image is well-composed with a clear focal point.”
  • Balanced framing: the subject sits well within the frame. “The balanced framing keeps attention on the subject.”
  • Strong focal point: attention is directed clearly. “The strong focal point immediately draws the eye.”
  • ×Awkward crop: important parts of the subject are cut off. “The awkward crop removes part of the subject’s arm.”
  • ×Distracting composition: unnecessary elements compete for attention. “The bright object in the background creates a distracting composition.”
Swap the vague for the specific
Instead of…Try…
“It looks weird.”Low contrast
“Something feels off.”Flat lighting
“Bad quality.”Soft focus
“It doesn’t pop.”Weak subject separation
“Too bright.”Overexposed highlights
“Washed out.”Undersaturated colors
There's something wrong with the body.
The cat in A is missing its back left paw, and its front right leg bends unnaturally backward.
The text looks weird.
A's sign reads "Cofffe," and the street name is a string of random symbols.
B has AI artifacts.
B shows six fingers on the right hand, and the wrist merges into the table.
The background is wrong.
A's crowd has three figures with near-identical faces: classic AI repetition.
Capitalize A / B Name the pixel-level detail Prioritize the loudest issue first Use the prompt's own wording

Do

  • +Zoom in inch by inch: most issues hide at normal size.
  • +Search when you need to verify a fact, flag, or count.
  • +Echo the prompt's own wording in your answer.
  • +Build a case: name specifics in both images.
  • +Stay consistent: your Overall Preference should match the issues you name.

Don't

  • ×Copy-paste responses across dimensions or tasks.
  • ×Use AI to help you answer.
  • ×Contradict yourself: a hand you flagged as mangled shouldn't win naturalness.
  • ×Settle for a tie: make the call whenever the images let you.
  • ×Write "A looks more polished" and stop there.
Your go-to reference before starting a new task type. Open the full guidelines for scoring dimensions, anchors, tips, and worked examples.

→ Open the full Task Guidelines document

Image
  • Text To Image Compare
  • Reference-to-Image Ranking
  • Search-Grounded Image Generation Eval
  • realism-hc-elo
  • Realism Quiz
  • Cv2 R2i Heldout
Video
  • Text-to-Video Ranking
  • Image-to-Video Ranking
  • Omni R2v Elo
Audio
  • Omni Tts Elo
  • Text-to-Audio - Audio-Only Evaluation
Multi-Modal
  • Multi-Capability Model Ranking
Ad Creative
  • Ad Creative Overall Ad Suitability (Q0)
  • Ad Creative-Visual Flaws (Q2)
  • Ad Creative – Photorealism (Q3)
  • Ad Creative– Visual Appeal (Q4)
  • Ad Creative – Text Flaws (Q5)
  • Ad Creative – Text Style (Q6)
Don't see your task type? Let your Pod Lead know so it can be added.

Overview

Everyone starts on image tasks. Once you qualify, you'll train on video and then get video comparison tasks mixed in with your image work. Here's how you become eligible, what training and calibration look like, and how to score each comparison.

How you qualify

You become eligible for video training once you’ve built a track record on image work: 15 active hours on multimango.com image tasks. Once you hit 15 image hours, you’re moved into video eval setup automatically.

What you’ll do

Qualifying is two steps:

  1. Video training walks you through the video task and the quality dimensions you’ll be judging. You’ll need at least 75% to move on to calibration, and you get two attempts on each knowledge check.
  2. Calibration: you compare two AI-generated video clips and judge them across the task’s 9 quality dimensions (video and audio), picking the better clip overall. You complete 4 calibration tasks, and 2 passing tasks qualifies you for video tasking.

Your time cap

You get 3 paid hours for training and calibration, and you complete your 4 calibration tasks within that time. Your timer warns you as you approach the cap, and time beyond 3 hours isn’t paid, so pace yourself. It’s built to fit inside the window.

While your Pod Lead reviews

You can’t continue tasking until you’ve completed all 4 calibration tasks, so get them done within your time cap. Once they’re in, you can go back to tasking on Multimango while you wait for your Pod Lead to grade them. Once you pass, video tasks unlock. To be clear, you don’t switch to video-only: you’ll see the same standard mix of image and video tasks you had on your dashboard before.

How to score: slightly, strongly, or tie

Each task shows you two AI-generated videos. For every dimension, decide which side wins and by how much — it’s a per-dimension call, made fresh on each line.

Slightly prefer

One video is better, but the gap is small. You can name the specific advantage — yet another careful rater could reasonably call it a tie or lean the other way.

Strongly prefer

One video is clearly better on that dimension and you’re convinced of it. Not a lean — a conviction you’d hold even if pushed back on.

Tie

The dimension genuinely doesn’t separate them: no relevant advantage, or strengths and weaknesses offset. The videos don’t need to be identical.

A strong pick doesn’t have to be flawless. Garbled on-screen text can still clearly beat heavy distortion. If your strong pick has a visible defect, name it and say why it still wins — if you can’t write that sentence, it’s a slight.

Slight vs. strong at a glance

What to considerSlightly preferStrongly prefer
Why it winsBetter, but by a fairly small margin.A clear, immediately visible advantage separates the two.
Winner’s flaws?Yes — and they may make the call closer.Maybe, but not enough to change the outcome.
The other video?Has problems, but balances them with strengths.Its problems widen the gap.
Room to disagree?Yes — a careful rater could call it a tie or flip it.Probably not — a careful rater should land in the same place.

One difference from image work: on video tasks, Overall Preference is your gut-feel call, unlike image tasks where it’s an in-depth final choice weighing all dimensions.

Common Calibration Errors

These are the most frequent issues identified in calibration and quality reviews:

  • Same pick every time: choosing the same video for every dimension instead of judging each one on its own.
  • Wrong dimension: writing about something that belongs under a different dimension, like describing visual quality under Instruction Following.
  • Rushed or generic: a justification that could apply to any pair of videos, with no specific detail or moment cited.
  • Contradiction: writing a rationale that argues for the video you didn't actually pick.
  • No explanation or cut off: leaving an explanation box blank, or submitting a justification mid-thought.

Project FAQ

Most asked
Filter by
A

Multimango

You'll complete at least 10 calibration tasks, labeled "[Calibration check]: Compare two images," with up to 20 attempts total. The 70% is based on the share of tasks you pass, not an overall star average, so aim to pass at least 7 out of 10. This 70% bar is for image calibration; initial training requires 75%, and Text-to-Video (video) calibration uses a 2-of-4 bar instead. These numbers can change, so check with your Pod Lead if you're unsure.

If you're not at 70% after your first 10 attempts, you'll be prompted to complete additional training and reach out to your Pod Lead: you'll still have attempts left at that point, so that's a great time to ask for support. If your first 5–10 calibration tasks aren't reaching a 70% passing grade, reach out to your Pod Lead for support right away so you don't risk being offboarded from the project. If you don't reach 70% within all 20 attempts, you'll be offboarded from the project. This is the Failed calibration reason.

Only after your Workada dashboard tells you to. Do not sign up on multimango.com on your own: you will not be approved unless you go through the dashboard flow and click the "I created my account" button.

  1. In your Workada dashboard, find the "Getting Started" card and click Set up on "Create Multimango account."
  2. In the pop-up, click Open multimango.com. Keep the Workada window open: you'll come back to it.
  3. On multimango.com, enter your workada.co email (shown on your dashboard when you click Set Up) and click Continue. Note, this should be workada.co, NOT workada.com!
  4. Create a password (at least 8 characters) and click Continue.
  5. Check your email for a 6-digit verification code, enter it, and click Continue.
  6. You'll see a "Welcome!" screen. Return to the Workada window and click "I created my account."
  7. Wait for approval (done manually: typically within 72 hours).
  8. Once approved, select "Text To Image Compare" as your task type, and you're working.

Always use your workada.co email, not a personal address. Accounts created with a personal email can't be approved.

Your code is sent to your workada.co address, then forwarded to the personal email you originally signed up with: check both "All Mail" and "Spam" in that inbox. If it's expired or never arrives, use the "Resend" button to get a new one.

Approvals happen manually, typically within 72 hours. Once you're approved and your account is active, you will get an email, and you'll see a "Tasks" option appear in the Multimango sidebar. There's nothing else you need to do while you wait.

Work whichever of these four task types is showing on your Multimango dashboard:

  • Text to Video
  • Multi-Capability Model Ranking
  • Text to Image Compare
  • Reference to Image Ranking

These can appear and disappear at random, so you may not see all four at once: but everyone should have at least one available at any given time. If your usual type isn't showing (or won't let you pick up a task), work one of the others that has available tasks rather than sitting idle.

Prioritise video and image tasks. Other task types sometimes appear on your dashboard, such as UD Caption or Caption Quality Ranking. Only work those if no video or image tasks are available to you.

Task availability for a given category can fluctuate. If one or more tasks have the message "No Active ELO Evaluations" when you attempt to access them, go back and try another category. Remember to refresh your screen to make sure no new tasks for our priority categories appear. Do not email support about receiving the No ELO message on Text to Image Compare. Contact your Pod Lead for guidance if you consistently see it across all other task categories, especially multiple categories.

Tasks can time out, and the type may temporarily drop off your dashboard afterward. Refresh and check the other available task types: it typically reappears. We're aware tasks sometimes time out sooner than they should and it's on our fix list.

Quality is also read from behavioral signals, like speeding through tasks, taking too long, or picking the same side (for example, all A) many times in a row. You'll usually get a warning before you're blocked, so treat one as a cue to slow down and refocus on accuracy.

This message can have more than one cause. It’s shown by Multimango, not by Workada, and the same wording appears in several different situations:

  • We moved you between Multimango annotator groups. When that happens, Multimango drops whatever task you had claimed at that moment and your submit fails. This has nothing to do with your work.
  • Multimango paused your account based on its own behavioral checks. Multimango has described these to us as patterns like moving through tasks unusually fast, leaving a task open for a very long time, or selecting the same side (for example, all A) repeatedly. Some of these pauses lift on their own after roughly 24 hours. Others don’t.
  • You’re in video calibration. Until your calibration tasks are graded, you can’t pick up image tasks, and this message is what you see when you try.

If you’ve seen this message, don’t assume it’s a verdict on your accuracy, and don’t assume you’re being offboarded. The pause lives entirely in Multimango’s system. It doesn’t appear anywhere in Workada, which means your Pod Lead and our Ops team genuinely can’t look up why a particular account was paused.

If you get the block message:

  1. Don’t retry the same task. It won’t go through.
  2. Refresh and claim a new task. For most people that clears it.
  3. Wait about 24 hours before assuming it’s permanent.
  4. If it follows you onto new tasks, or no tasks appear at all, message your Pod Lead with your email, the task type, and roughly when it started. Please go through your Pod Lead rather than the support email for this one, so we can tell an account-level pause apart from a group change.

Let your Pod Lead know. They can review it and escalate it for you if the QA looks incorrect.

The only locations you can work from are the US and Canada, plus US territories by default (PR, GU, VI, AS, MP).

Review the dimensions first, then split your answers up by category, calling out different inaccuracies for each dimension instead of repeating the same point.

Instruction Following asks “did the image do what the prompt asked?” You’re checking the prompt’s requirements: subjects, objects, and actions; counts; attributes like color and clothing; style and setting; and negative constraints like “no text” or “sun not visible.”

Correctness asks “are the checkable facts and structures right?”, whether or not the prompt asked for them: spelling and legible text, accurate counts, anatomy and physics (limbs, shadows, reflections), and factual labels, logos, and flags.

Rule of thumb: Instruction Following is prompt faithfulness; Correctness is verifiable errors.

Move on to the next task. Once a task is closed you can’t go back to the same prompt.

Write a sentence for every applicable dimension, on every task type. Some tasks show all the dimensions to judge individually with a single justification box at the bottom: even then, cover each dimension you scored rather than writing one general comment.

It also helps a QA reviewer understand why you preferred one response over the other.

Refresh the page to get a new task. Don’t post screenshots of the material in community channels: sharing it exposes other Contributors to the same content. We share reported examples with the Multimango team and are pushing for better filtering on their end, but unfortunately can’t make any guarantees.

No, these don’t need to be flagged to the MM Task Issues channel.

B

Tasks and workflow

Tasks are available through the Task tab on your Dashboard and/or on Multimango (see next section for more details). Once your onboarding session is complete, your training is passed, and your account is fully set up, you'll be able to browse and pick up available tasks.

StatusWhat it means
Not startedThe task has been assigned or made available, but work hasn't begun.
In progressYou've picked up the task and are actively working on it.
You've completed and submitted the task for review.
Pending reviewYour submission is queued with the quality team: this is normal. No action needed unless it's been over 48 hours.
ReviewedThe task has been reviewed. You may receive feedback or a score.

There's no strict time limit, but aim for around 12–17 minutes on Text To Image Compare tasks and around 20 minutes on video tasks in the Multimango project. The 12–17 minute range applies to Text To Image Compare specifically: other task types, such as Reference to Image Ranking or the caption task types, routinely take longer and aren't held to it. Tasks on multimango.com are capped at 20 minutes, and image and video calibration tasks at 30 minutes, so going over these can affect your metrics. Other projects may differ. Trust your gut on the obvious calls, zoom in where it matters, and only fact-check claims that are actually checkable: a count, a label, a spelling. Don't go down rabbit holes.

Quality is what matters most, so take the time you need to do the work well. As you build experience, you'll naturally get faster without losing that quality.

Once you pass training, you'll need to complete calibration tasks first: see the Multimango section below for details.

Yes, either way works.

There isn't a rule that fits every task, so use your judgment. Slightly is the default, but don't hold back strongly when one option clearly wins on that dimension and you can defend it.

Refresh the page to load a new prompt.

Long prompts on image and video tasks can eat your whole handle time. Work them in this order:

  1. Review the outputs on their own, before you read the prompt.
  2. Read the prompt and identify the primary prompt: what the output basically has to be.
  3. Skim the rest, separating the details that matter from the minor ones.
  4. Make your selection on that basis.
  5. Write your justification against the details you judged important.

If a prompt is so long that this still isn't workable, let your Pod Lead know.

Take the time the task actually needs. Where it is unreasonable to complete a task properly inside the limit for that type, quality comes first.

You won't be blocked for going long while you're working in good faith. That only becomes a problem when the time is far above average and starts to look like time milking.

Yes, as long as your Workada timer still shows MultiMango as your active tab. Don't use one if it significantly increases your handle time for little or no gain in quality.

C

Offboarding

Yes, you may apply if you’d like. Keep in mind that the Data Labeling Specialist role has the easiest training and calibration, so other roles may be more challenging.

Everything for the Sheets project: your Slack channels and the questions we hear most.

Channels

Sheets FAQ

A

Getting started and setup

Welcome aboard. Here’s your path from here:

  1. Verify your identity.
  2. Set up Slack.
  3. Set up your Workada Timer.
  4. Accept the terms and conditions.
  5. Complete the training modules.
  6. Complete the calibration task.
  7. Complete your first two tasks.

Your first two tasks go to a reviewer to check and accept. Once they pass, you’re cleared to claim live tasks from your Dashboard.

Most Contributors finish in under 2 hours, and we compensate up to 2 hours of training time. If you need a little longer to feel ready, that’s completely fine. Anything past the 2-hour mark just isn’t compensated.

Almost there. Your first two tasks go to a reviewer first. Once they’re accepted, you’re cleared to claim live tasks from the main pool on your Dashboard.

Yes.

Answer coming soon.

B

Tasks and reviews

Hang tight, this is normal. Your first tasks are assigned to a reviewer by hand, so they can take longer than pool tasks. If yours has been waiting more than 5 business days, reach out to your Pod Lead and they’ll help move it along.

We’re actively working to improve these metrics. For now, here’s what we capture for each task:

  • Claude score on first attempt. The review rating your task gets the first time you submit it, as it moves from edit_task into task_review. The review system scans the prompt and workbooks for factual contradictions and critical coverage gaps, which makes it the cleanest read on how close your task is on the first try. A score of 3 or below sends the task back for a redo; a 4 or higher, with the rubric landing around 95 to 100%, gets it accepted.
  • Distribution of Claude score on first attempt. The share of all submitted tasks that hit a 4 or 5 on the first attempt, our first-pass quality rate. It shows how consistently work lands right the first time instead of leaning on redo cycles. Goal: as high as possible.
  • Redo rate. How many times a task gets sent back to redo_task before it moves through to audit. A task is redone when the review rating comes back at 3 or below, or the rubric eval flags items that may be wrong. Most tasks should settle within one iteration. Goal: as few as possible, ideally zero.
  • Average handling time (AHT). The time between claiming a task from the pool and submitting it. Most tasks should take around 45 minutes, including any revisions, though this varies with complexity.

We aim for around 45 minutes per task on average, including any revisions and redos. Some tasks run longer and some run shorter depending on complexity. If a complex task needs closer to an hour to finish properly, that’s fine. Consistently running well over an hour is what we’ll follow up on.

New tasks drop in batches throughout the day, at no set time, so keep an eye on your Dashboard and claim them as they land. You can hold one task at a time. If you claim a task and don’t work on it within 24 hours, it detaches and returns to the pool for someone else, so only claim a task when you’re ready to start it.

To be confirmed.

Keep iterating first, since most tasks clear once you’ve addressed the real issues. But if the only flags left are hallucinations or oversensitive calls you disagree with, and the task is otherwise a solid 4 or 5, submit using the “Submit with Eval Override Request.” Add a few short bullets on why you think the flag is wrong.

Use this only when you truly believe a flag is incorrect, not as a shortcut past valid feedback.

C

Working location

Yes, you can work from any country.

Get oriented, get connected, and start building chart tasks that a frontier model gets wrong.

Channels

About the project

What Chartography is

Today’s frontier AI models read a simple bar chart without trouble. They fall apart on the charts professionals actually make decisions from: Kaplan-Meier survival curves, candlestick charts, contour maps, Sankey diagrams, Bode plots, phase diagrams, control charts, and weather soundings.

Chartography exists to find and document those failures. Each task pairs one complex chart with one hard question, one verified answer, and a clear step-by-step solution. Every task also has to clear the complexity bar, which means a frontier model has to get it wrong.

A brilliant question with a fuzzy answer isn’t usable. A rock-solid answer to a question the model already nails isn’t usable either. All five pieces together are what make a task deliverable.

What makes a strong chart task

  • A complex visualization that rewards careful reading: multiple axes, overlapping series, dense or reused legend colors, a technical domain.
  • A question with one definitive, checkable answer that takes several reasoning steps to reach.
  • A clean step-by-step solution a reviewer could follow and independently verify.
  • A model failure, documented, not asserted.

The bar: the model has to fail

Every task you create runs against a frontier model three independent times. A task is at an acceptable complexity level if two or more of those three attempts score as failures. If the model gets it right, the task isn’t hard enough yet, and that’s not a dead end: it’s your signal to tighten the question and try again.

Iterating on your own prompt until the model breaks is the core skill of this project, and it’s the part we most want your feedback on.

The domain field

Every task is tagged with a domain. Pick the closest fit; use Other if nothing matches.

  • Finance and economics: Finance, Investing, Economics
  • Science: General STEM, Chemistry, Biology, Physics, Geosciences
  • Applied: Manufacturing, Supply Chain, Healthcare, Other
  • Engineering: General Engineering, Mechanical, Electrical, Civil, Environmental

How to Build a Strong Task

A complex chart does a lot of the work for you, which makes this the place to spend time before you write anything. Whether you’re bringing your own chart or choosing one from the Chart Library, look for something busy and difficult to parse at first glance.

Strong charts usually have several of these features:

  • Large numbers of overlapping data points
  • Multiple axes with different units and scales
  • Legends with many elements, especially ones that draw a distinction between similar colors or shapes
  • Different types of data displayed together on the same plot
  • Many clearly enumerated subdivisions
  • Multiple curves that intersect at various points
  • Colored areas that overlap in irregular ways

The list isn’t exhaustive, so if a chart makes you slow down and look twice, it’s probably a good candidate.

Neighborhood environment interventions map with overlapping data points in three colors.
A dense spatial map with three overlapping point types and a boundary outline.
USDA Crop Progress chart with three stacked panels and multiple years of overlapping data.
A multi-panel crop report with stacked areas, five years of trend lines, and growth-stage curves.
Grid of 14 convergence plots comparing four algorithms across different functions.
A grid of 14 convergence plots comparing four algorithms, each with different scales.

Understand the chart before you write. You either know what the chart shows out of the gate, or you spend time with the chart description and source URL to get up to speed. Either way, knowing what the data actually shows is what lets you write an informed, realistic prompt that stays unambiguous and points to a single defensible answer.

This guide covers what a strong prompt has to do, the seven guiding principles to check before you submit, and the question types that make the model reason instead of just count.

Read the Question Types & Prompt-Writing Guide before you write your next task, and keep it open while you work.

Prompt realism

The prompt should ask about the data underneath, not visual elements. Read the chart description in the claim sheet and open the source URL before you write. Then sense check: would a researcher ask this about this chart, and would the answer be a valuable insight in a professional setting?

Defensible, verifiable answers

If the scope allows several readings, it means several viable golden answers, and that’s a problem. Build the question on something that can’t be misinterpreted and layer guardrails into the prompt. Then sense check: would two reviewers following your solution land in the same place, and is anything about your golden answer debatable?

Every task you submit is reviewed against the quality rubric. Your reviewers work from that same document, so read it once before you start tasking and keep it open while you work.

If anything in this guide seems to disagree with the rubric, the rubric is the standard you’re graded on, so flag it and we’ll fix the guide.

  • The chart made you slow down and look twice.
  • One defensible answer, and you can say why.
  • The answer comes from the image, not from general knowledge or a printed label.
  • Scope names the region, the series, and the form of the answer.
  • Your solution walks a reviewer to the answer, including the parts that are easy to miss.
  • A researcher could plausibly ask your question.

Best ways to source your own chart for task acceptance

Chart choice shapes everything that comes after it. A chart that’s too simple limits how complex your prompt can get, no matter how much time you spend on it. This guide covers what to look for, where to look, and how to move from chart to prompt.

Seeded charts are available on your dashboard, and you can build strong prompts from those alone. Sourcing your own chart isn’t required.

That said, learning to source well gives you an option worth having. It lets you be proactive when tasking and bring in charts you already understand, in domains where you know the terminology and can spot the ambiguity faster.

The strongest charts share a few traits: multiple variables, overlapping series, or data points without clear labels, etc.. That ambiguity is what gives you room to build toward the more demanding question types instead of settling for a simple value extraction.

A chart that’s visually clean and easy to read at a glance usually works against you. There’s not enough there to construct a prompt that’s hard for the model to answer, and you’ll end up spending more time forcing complexity than the chart can support. If a chart takes you less than a minute to fully understand, it’s worth moving on.

Examples of charts:

→ this chart above is NOT a good sourced chart as it is too simple
This chart above is a good sourced chart as it has ambiguity (different colors, overlapping part)

Searching directly for charts, rather than starting from an article and hoping it contains one, tends to work best. You can keep a few charts in rotation at a time. If one isn’t yielding a strong task, you can switch to another instead of getting stuck reworking the task on a chart that AI is easily understanding.

Complexity only helps if the chart is actually legible. Before you settle on one, make sure the image is high enough resolution that axis labels, legends, and data points are clearly readable, not blurry, cropped, or compressed to the point of pixelation. Watch for watermarks or overlays that obscure part of the chart, and avoid screenshots where text has become distorted or cut off at the edges.

A chart that represents genuine data complexity is what you want. A chart that’s ambiguous because you can’t actually read it clearly is a different problem, and it will cause issues at review regardless of how strong the prompt is.

Before settling on a chart, read the article or the relevant section of the paper it comes from. This tells you whether there’s enough domain depth to support the question types that lean on interpretation or reasoning, and it surfaces the terminology you’ll want to use in your prompt. Charts backed by a source with real analytical depth consistently produce stronger prompts than charts pulled with no context.

Choose your question type after you’ve studied the chart, not before. Starting with “I want to write a calculation question” and then hunting for a chart that fits it usually leads to a forced prompt. Look at what the chart actually offers, and match it to one of the six question types:

  • Value extraction: the chart lets you read a single value or range directly at a specified point, with little to no transformation needed.
  • Comparison and ranking: the chart supports comparing multiple values, ordering them, or identifying an extremum, without needing a calculation to get there first.
  • Calculation: the chart gives you clean numerical relationships to work with, so you can build a self-contained computation, a difference, ratio, growth rate, or similar.
  • Enumeration: the chart has a clearly bounded set of items you can count or list against a stated condition, with unambiguous item boundaries. (Now Paused)
  • Multi-step reasoning: the chart supports chaining two or more different kinds of operations together, where one step’s output feeds a different kind of step, like deriving a value and then comparing it.
  • Conceptual interpretation: the chart requires understanding what it represents, not just reading it, whether that’s a qualitative judgment, identifying something from a signature, or locating a meaningful feature.

A chart with real complexity often supports more than one of these. Picking the type the chart is strongest for, rather than forcing a type it doesn’t support, is what separates a well-sourced chart from one that fights you.

  • Does the chart have overlapping series, unlabeled points, or multiple variables?
  • Would understanding it fully take you more than a minute?
  • Does the source article or paper add real depth you can draw on?
  • Have you matched the question type to what the chart actually supports, rather than the other way around?

If you can say yes to all four, you’re working from a chart that gives your prompt room to succeed.

Better ways to stump SOTA

Stumping the model is not about making the language of the prompt confusing or relying on clever wording. Rather, the difficulty should come from the reasoning required to solve the task. The strongest prompts contain a genuinely analytical step grounded in the chart, the domain, or both. That might involve interpreting an ambiguous pattern, connecting multiple variables, applying domain knowledge, performing a multi-step calculation, or deciding between competing conclusions using evidence from the chart.

Contributors with strong SOTA pass rates are generally not making their prompts harder to read. They are choosing charts and questions where the underlying reasoning is genuinely difficult, while keeping the prompt itself specific, detailed, and answerable.

Precision on the specific, not the general. The strongest prompts are often narrowly scoped. Instead of asking about the chart as a whole, direct the model to a specific part of the figure where careful reasoning is required. This could be a particular y-axis threshold, a defined region of the parameter space, a specific panel, or the behavior of a set of lines within a limited window.

Keep prompts short whenever possible, ideally one or two questions. Define the panel, time or value range, boundary rule, and whether the bounds are inclusive, then stop. Avoid adding extra layers that do not contribute to the core reasoning.

A strong model-stumping task usually starts with choosing the right chart. Focus on what the chart makes the viewer work to distinguish, not simply on the subject matter it covers.

Charts with overlapping series, unlabeled points, closely spaced values, multiple variables, crowded legends, or competing visual patterns tend to support harder questions than clean charts with a few clearly separated lines.

Before writing the prompt, spend a few minutes identifying where the chart itself requires careful interpretation. Look for regions where series converge or cross, values are difficult to distinguish, multiple conditions must be considered at once, or the same observation can be interpreted differently depending on context.

The goal is not to exploit poor readability, but to find a part of the figure where careful visual reasoning and domain understanding are genuinely required.

Beyond the general principle of finding ambiguity, these are seven recurring patterns worth checking for when picking a chart and a question.

Pattern 1: The sliver. The answer depends on something tiny: a thin line inside a bar, a near-zero bar next to tall ones, a 1% slice, a hairline ribbon, a tiny squeezed region. The model reads coarsely and misses anything a few pixels wide, even when a "none" or "zero" answer is correct. The model would rather find something than report nothing.

  • Look for: a thin line in a bar, a near-zero bar, a 1% slice, a hairline ribbon, a tiny region
  • Ask: count, which one, or is there any, so that missing the sliver changes the answer

Pattern 2: Shade discrimination. Two similar colours, and the answer depends on telling them apart: several steps of the same hue, adjacent similar hues, or a legend with more entries than you can hold in your head. The model lands one step off on a colour ramp and confidently matches the wrong legend entry. The colour is the category here, not decoration.

  • Look for: 6 or more legend colours, a colour ramp with 4 or more steps, similar neighbouring hues
  • Ask: which category or which band, named exactly as the key names it

Pattern 3: The decoy. The right answer isn't where your eye goes first. The model goes for whatever is biggest, steepest, or most obvious, so if the correct answer sits in a squashed or off-to-the-side part of the chart while something large and obvious is wrong, it takes the bait. Reading conventions make great decoys: box top vs. whisker end, local vs. global peak, a log or flipped axis.

  • Look for: a squashed region near an axis, local vs. global maximum, box top vs. whisker, a log or inverted axis
  • Ask: largest, first, or highest, defined precisely, where the obvious answer is wrong

Pattern 4: Two-key filter. Each condition is easy alone. Both at once is not. Colour and shape, series and panel, a threshold on x and on y: holding two visual conditions across many candidates makes the model drop items that qualify and admit ones that don't. Per-panel legends are especially effective, since the model reads the wrong panel's key.

  • Look for: colour and shape legends, per-panel legends, thresholds on both axes
  • Ask: which items meet both conditions, with a small answer set

Pattern 5: Near-ties and hairline crossings. The model rounds small differences away: two bars almost the same height, a curve poking a hair above another, two close crossings. This has become the most popular pattern, which is a problem: on multi-line charts it can turn into a tolerance argument rather than a real stump. Use it only where the gap is genuinely visible.

  • Look for: almost-equal bars, hairline crossings, two crossings in a narrow window
  • Ask: strictly higher or lower within a stated window, or how many crossings between a and b

Pattern 6: Boundary gap. Ask for the width of a band, not the position of its edge: the thickness of a stacked slice or ring, not a drawn boundary. The chart never prints that number, so the model reads the edge, or the running total, instead of the individual gap. The bands can be perfectly ordinary in size. This isn't about something small. It's about measuring the wrong thing.

  • Look for: stacked areas or bars, ribbon or band charts, radial stacks, anything drawn as a thickness
  • Ask: how thick is this band at x, where the outer edge peaks somewhere else entirely

Pattern 7: Occlusion recovery. The data is there, just hidden behind something else: a marker or line segment covered by another mark. Distinct from Pattern 1 (small but visible) and Pattern 6 (genuinely absent), this is plotted but hidden, and the answer has to be reconstructed from what's on either side. Never tell the model the value is hidden. That collapses it into an ordinary read.

  • Look for: crowded time series with overlapping lines, stacked scatter markers, later series drawn over earlier ones
  • Ask: what is X's value at this point, where X is buried under other series there

For all seven: the pattern tells you which chart to pick and where the hard part is. It doesn't change how you word the question. That still has to be the ordinary thing a professional in the field would ask, in the chart's own vocabulary. A question that only makes sense as a way of catching the model out will be sent back as unrealistic.

Before working on the prompt, take time to review the source URL from which the chart was derived. This will give you the broader context needed to understand what the chart is measuring, what each variable represents, and what relationships or trends the figure is intended to highlight.

Reviewing the source also helps you use the correct domain-specific terminology and frame the task in a way that reflects how a professional in that field would reasonably interpret the figure. That context should inform the prompt, Golden Answer, and Step-by-Step solution so that all three are conceptually accurate, technically precise, and grounded in the source material rather than relying only on surface-level visual reading.

Ask the question a professional would ask. Focus on three things: what the underlying data represents, how the figure contributes to the broader analysis or argument, and what someone working in that field would realistically want to learn from it.

Use that context to frame the prompt around a meaningful analytical objective rather than a surface-level visual task. The question should reflect how a domain professional would interpret, compare, or use the data, while the Golden Answer and Step-by-Step solution should apply the same terminology and reasoning consistently.

Of the six question types outlined in the Chartography Question Types and Writing Guide, prompts that combine conceptual interpretation with multi-step reasoning are generally the most effective at challenging the model. By contrast, simple value extraction is usually the easiest question type for the model to answer correctly.

If your prompts are passing too easily, check whether they are asking the model only to read values from the chart rather than reason with them. Stronger prompts should require the model to connect multiple observations, perform calculations or comparisons, and interpret what those results mean in the context of the chart.

Reaching that level of depth often requires going beyond the figure itself. Read the source article or paper, understand what the variables represent, and spend a few minutes looking up unfamiliar terminology. That context makes it much easier to create a question that reflects genuine domain reasoning rather than surface-level chart reading.

Difficulty should come from the reasoning required, not from complicated wording. Ask a question that a professional in the field could reasonably ask, and make sure two careful reviewers looking at the same evidence would arrive at the same answer.

Sometimes a prompt fails to challenge the model because the chart itself does not contain enough complexity to support a difficult, meaningful question. Precise wording cannot fully compensate for a figure with limited data, relationships, or analytical depth.

If you have already tried clearly defining the scope, using appropriate domain-specific terminology, and combining conceptual interpretation with multi-step reasoning, but the task still passes too easily, the chart itself may be the limiting factor.

At that point, avoid forcing complexity through convoluted wording, arbitrary conditions, or unnecessary question chaining. It is usually more effective to select a different chart that naturally supports deeper comparison, calculation, and interpretation than to continue reworking a fundamentally simple figure.

  • Have I reviewed the source URL and understood what the chart is actually measuring, why it matters, and how someone in the domain would use it?
  • Would a professional in this field reasonably ask this question, or does it feel like a visual scavenger hunt created only to stump the model?
  • Am I using the chart's domain-specific terminology rather than generic language?
  • Does the task require conceptual interpretation, calculation, comparison, or multi-step reasoning rather than simply reading a value, colour, label, or count?
  • Does the chart itself contain enough complexity to support the question, such as overlapping series, unlabeled observations, multiple variables, crossings, thresholds, or competing patterns?
  • Am I using the chart's underlying data as part of the reasoning rather than asking only about visual characteristics?
  • Is the prompt short and focused, ideally one or two connected questions, rather than a chain of unrelated tasks?
  • Does each question build naturally on the previous one, rather than question hopping between unrelated parts of the figure?

Examples of Stellar Work

01
Model-stumping
Include at least one element designed to trip up an AI that isn’t reading carefully.
02
Defensible
Every golden answer should be verifiable by anyone studying the chart closely.
03
Step-by-step solution
Your solution walkthrough should leave no ambiguity about how to arrive at the answer.

Browse examples by domain: click a tab below to switch.

01 · Economics
Eurostat HICP inflation chart, Jul 2016–Jul 2026
Eurostat HICP — Annual rate of change, Jul 2016–Jul 2026
PromptRank the following CPI components from highest to lowest based on their inflation rate in July 2020. Then, identify which component exhibited the lowest absolute deviation from 0% inflation over the period from July 2016 to July 2026, indicating the component whose inflation rate remained closest to price stability throughout the period.
Golden output
In July 2020, the components ranked from highest to lowest based on their annual inflation rate were:
1. Food, alcohol, & tobacco
2. Non-energy industrial goods
3. Services
4. All-items
5. Energy

From July 2016 to July 2026, Non-energy industrial goods remained relatively closest to 0% inflation, indicating the most stable inflation rate around the zero-inflation benchmark among the components listed.
Step-by-step solution
To arrive at the correct solution, first refer to the legend at the bottom center of the graph to identify the five components tracked: All-items, Food, alcohol & tobacco, Energy, Non-energy industrial goods, and Services. Next, use the x-axis to identify the timeframe covered by the graph, which runs from July 2016 to July 2026. Locate July 2020 on the x-axis, then trace upward to the corresponding data points for each component. Use the y-axis, which measures annual inflation percentage, to compare the five components and rank them from highest to lowest inflation rate. The rankings for July 2020 are: 1. Food, alcohol, & tobacco, 2. Non-energy industrial goods, 3. Services, 4. All-items, 5. Energy. Finally, to determine which component remained closest to 0% inflation between July 2016 and July 2026, examine each component's trend across the entire timeframe. Starting from July 2016 on the left side of the graph, track each line and compare how closely it remains to the 0% inflation level. Non-energy industrial goods remained relatively closest to 0% throughout the period, making it the component with the most stable inflation rate around the zero-inflation benchmark.
Key lessonA two-part question that requires both a snapshot read and a longitudinal analysis is harder to fake and produces a richer golden answer.

02 · Finance · Investing
Index price chart: SPY, EEM, TLT, COY, GSP, RWR (2010–2013)
Index prices — SPY, EEM, TLT, COY, GSP, RWR (2010–2013)
PromptFrom the provided graph, what are the second and third highest priced indexes in January 1st 2011 and January 1st 2012? What is the approximate rate of return of COY index between January 1st 2010 and January 1st 2013?
Golden output
January 1st 2011: second index is EEM, and third index is SPY.
January 1st 2012: second index is TLT, and third index is SPY.

There is no information to answer the question.
Step-by-step solution
To answer the question, we first need to identify the different index price lines plotted on the graph. To do this we look for the legend with the different color series at the top left section of the chart. Here we can identify the six indexes represented with six different colored lines: index SPY is represented with a black line, index EEM is represented with a red line, index TLT is represented with a green line, index COY is represented with a blue line, index GSP is represented with a cyan line, and index RWR is represented with a purple line.

Next, we need to identify the horizontal axis on the bottom part of the chart, which is labeled “Time”, and that has markers in the following points in time: from left to right, January 1st 2010, January 1st 2011, January 1st 2012 and January 1st 2013.

Having identified the different points in time, we now focus on the two dates that are specified in the prompt: January 1st 2011 and January 1st 2012. Starting with the first date, we need to check the relative vertical position of each index price line at that point over the horizontal axis. Maintaining our position over that point in time, we check the position of each index price line from top to bottom. The indexes appear as follows: RWR, EEM, SPY, GSP, TLT.

Moving on to the second date, we repeat the process, checking the relative vertical position of each index price line at that point over the horizontal axis. The indexes appear as follows: RWR, TLT, SPY, GSP, EEM.

Having checked the relative position of all the index price lines at both points in time, we can now answer the question: for January 1st 2011, the second index is EEM, and third index is SPY; for January 1st 2012, second index is TLT, and third index is SPY.

To answer the third question we need to locate the COY price index line at both points in time. However, we can see that the chart does not show any data points for COY index. Therefore, we can conclude that there is no information to answer the question.
Key lessonIncluding a legend entry with no matching data line is one of the most reliable ways to expose hallucination. The golden answer needs to explicitly call out the missing data, not estimate it.
01 · General STEM
Optical density vs. wavelength chart, curves 1–7
Optical density (−log T) vs. wavelength (μ) — curves 1–7
PromptHow many times does curve 5 intersect with curve 6 between the wavelengths of 0.7 and 0.8 microns?
Golden output
They intersect exactly 2 times.
Step-by-step solution
Locate curves 5 and 6 identified by their labels on the leftmost region of the chart. Trace both curves toward the 0.7 - 0.8 range on the x-axis. Within the interval, count the number of points where the two curves cross paths.
Key lessonSimple does not always translate into easy. A one-line prompt with clear units and bounded scope can be harder to answer correctly than a long one, and leaves the golden answer completely defensible.

02 · Geosciences
Chart
PromptInto how many eons is the chart divided into? Per this chart, for how many of those eons does multicellular life overlap and what are the names of those eons? Which of the reported glaciations lasted the longest?
Golden output
The chart is divided into 4 eons. Multicellular life exists for 2 of those eons, the Proterozoic and the Phanerozoic. The Pongola glaciation.
Step-by-step solution
The outermost end of the spiral serves as a legend for what each stratum of the spiral represents. Eons are labeled as the innermost stratum of the spiral. If we follow the innermost stratum all the way towards the center of the spiral, we will notice that it is divided into 4 distinct colors, each labeled with a different name (starting from the inside: hadean, archean, proterozoic, and phanerozoic). This is how many eons the chart is divided into.

Going back to the outermost end of the spiral, we can notice that the thin stripes making up the outermost strata of the spiral represent the presence of different organisms, each one labeled accordingly with a different color. Multicellular life is the 5th band from the outside going in, and it is colored in a forest green shade. If we trace this line back to the inside of the spiral, we can notice that it overlaps with two different eons, the phanerozoic and the proterozoic.

Notice that glaciations are represented as blue segments within the innermost tier of the spiral. Since the spiral represents a linear timeline, the arc length of each segment is directly proportional to the length of the event it represents. Therefore, the longest glaciation is represented by the longest blue glaciation segment, which is clearly the pongola glaciation.
Key lessonThree-part question, clear & defensible answers, clear SOTA failure on one of the questions

03 · Chemistry
Chart
PromptIs there a peak above 1 total ion currents on this chromatogram that is not labeled with a number? If so, how many?
Golden output
9 total peaks above 1 total ion current not labeled with a number on this graph.
Step-by-step solution
Scan the chromatogram for any spike lacking a dashed leader line and a printed number, the same feature every other peak displays. Then scan each spike in relation to the left side gridlines of total ion currents above 1. There are two peaks that are slightly below the 1 total ion current line near the 37:00 mark which should not be counted leaving total peaks not labeled with a number at 9.
Key lessonNear-boundary cases are where precision counts. Prompts that include items just below a threshold give you a clear way to separate careful readers from rough estimators.

04 · Biology
Chart
PromptIn panel A, the group Singapore is assigned a specific color in the Group legend. Find that color; we will refer to it as color A. Use the legend to identify the site types Singapore has. How many of each site? Next, locate Singapore on the horizontal axis of Panel B. For the Singaporean site that has the LEAST quantity from the sites you counted previously, what is its color in Panel B?
Golden output
1 LR Market site, 8 field sites. The LR-market site is colored a shade of green.
Step-by-step solution
Look at the "group" legend on the right side of panel A. The group "Singapore" is represented by dark purple color.

Determine the site type; an LR-Market is a triangle. Look for dark purple triangles. There is exactly one purple triangle on the plot on the left. There are 8 field sites represented by dark purple dots. There is only 1 LR market, so we have to find what color the LR market is attributed to in panel B. It should be green with the caption "unknown".
Key lessonMulti-hop questions where each step depends on the previous one expose reasoning gaps that single-lookup questions can’t. The model must carry information across visual elements and panels.

05 · Physics
Chart
PromptHow many times does the plot with spectral width 0.1nm intersect the y=50% transmission line? Rounding to the nearest major gridlines, give me the tightest possible upper and lower bounds that will cover the entire range of transmission values for each spectral width. Return each answer as an ordered pair where the first value is the lower bound and the second value is the upper bound. Assume that the gridlines themselves are INCLUSIVE of the range.
Golden output
The plot with a spectral width of 0.1nm intersects the 50% transmission line 20 times.

Spectral Width = 0nm ; (0,60)
Spectral Width = 0.1nm ; (20,60)
Spectral Width = 0.3nm ; (30,60)
Spectral Width = 0.5nm ; (40,50)
Step-by-step solution
Question 1: notice that the plot corresponding to a spectral width of 0.1nm intersects the horizontal dotted line y=50% transmission twenty times.

Question 2: For each spectral width chart, find the smallest window of Y-values that will include all the graphed data. You are only allowed to choose values of Y that are labeled on the Major gridlines however (multiples of 10, from 0 to 100), so for example if the lowest data point is 15%, the lower bound of the data will have to be 10%. Similarly, if the highest value is 55%, the top bound will have to be 60%. The fact that the bounds are inclusive means that they include their own values in the reported range, e.g., if the curve's maximum is touching but not crossing 50%, then 50% would be an appropriate upper bound.
Key lessonDefining boundary rules explicitly in the prompt (“inclusive”) turns an ambiguous edge case into a hard test. Models that skip the rule produce wrong answers; your golden answer proves it.
01 · Healthcare
Chart
PromptBetween 1800 and 2024, how many distinct times does France’s child mortality rate reach or exceed 30%? If France’s rate remains at or above 30% for multiple consecutive years, count that continuous period as one occurrence.
Golden output
Between 1800 and 2024, France’s child mortality rate reaches or exceeds 30% on 8 distinct occasions.
Step-by-step solution
To arrive at the correct solution, first refer to the legend in the bottom-right corner of the graph to identify the countries being tracked and the color assigned to each country. The legend shows France as light orange. Next, use the x-axis, which represents years, to identify the timeframe being analyzed: 1800 to 2024. Then, focus specifically on France’s light orange line. Use the y-axis, which measures child mortality rate as a percentage, to identify the 30% threshold. The dashed horizontal reference lines can be used as a guide to determine whether France’s light orange line reaches or rises above 30%. Starting at 1800 on the left side of the graph, follow France’s light orange line chronologically from left to right until reaching 2024 on the far right side. As you follow the line, count each distinct occurrence where France’s child mortality rate reaches or rises above the 30% threshold. Do not count consecutive years separately if the line remains at or above 30%; instead, count each distinct occurrence of reaching or crossing the threshold. After reviewing the entire 1800–2024 timeframe, France’s child mortality rate reaches or exceeds 30% a total of 8 times.
Key lessonThe scoping of the prompt is really clear which makes the golden output defensible. When a chart spans a long time period, counting rules that define what “one occurrence” means are what make the golden answer defensible against any reasonable reading.
01 · Mechanical Engineering
Chart
PromptHow many types of geomaterials are neutral, partially or wholly falling within 5-10 Ks (MPa/mm), 0.1-10 tp-tr (MPa), and 1-10e-2 (tp-tr)^2/Ks?
Golden output
There are 5 types of geomaterials that fall within the neutral range outlined: limestone, sandstone, marble, granite, and serpentinite joints.
Step-by-step solution
When looking at the X [Ks (MPa/mm)], Y [tp-tr (MPa)], and Z [(tp-tr)^2/Ks] axes in the ranges identified, create an outline to define the region and count the areas that are within that region of the Ashby plot. Aqua, green, red, yellow, & blue areas are within the region and are labelled with the specific joint types.
Key lessonThree-axis Ashby plots are a rich source of tasks because they require simultaneous reading across multiple dimensions. Partial-overlap counting language (“partially or wholly”) makes the rule explicit.

02 · Civil & Environmental Engineering
Chart
PromptHow many groundwater samples contain ion concentrations and ratios indicative of silicate weathering? Count samples even if they only touch the silicate weathering boundary.
Golden output
7 markers.
Step-by-step solution
A marker on this chart will be identifiable as a black square. Locate the box in the middle of the chart that says "Silicate Weathering." Count the number of black squares that are inside or touch any part of this box. There are seven. Three squares are overlapped, and may appear as a larger shape, but they are three separate markers.
Key lessonOverlapping markers are a powerful precision test. State the boundary rule explicitly and your golden answer is bulletproof, anyone who counts carefully gets the same number.

Chartography FAQ

A

Training and Calibration

Training and calibration have a combined paid time cap of 1.5 hours for Chartography. This limit has been set to cover everything you need to complete both and we're confident it provides enough time for most contributors. Once you reach the cap, you're welcome to continue with your timer turned off, but any time beyond that won't be compensated.

Your Pod Lead reviews your training and you are notified once they've gone over it. Once you pass, you move on to calibration.

Calibration is a task you complete once you've passed training. You get two attempts to pass it. If your first attempt doesn't pass, it comes back to you to revise and try again. Your Pod Lead reviews your calibration task, so if you've been waiting a while, check in with them for the current timeline in your pod.

B

Writing strong tasks

It means a frontier model can already answer your question reliably, so the task isn't usable as is. The bar in Chartography is that the model has to fail. If a task comes back "Within SOTA capability," iterate to make it harder: more multi-step reasoning, more disambiguation, fewer direct reads.

This one trips people up because it sounds negative, but it's good news. It means the model failed to answer your question, which is the goal. If a task instead comes back as "Within SOTA capability" (you may also see "previous attempt" or "draft"), that means your question wasn't strong enough yet.

Skip a chart when you can't make it both single-answer/verifiable and hard enough, even after trying the usual difficulty levers: multi-hop reasoning, miscounting traps, overlaps, legend or color traps, scale changes, absence answers, off-chart label reading. Mark a chart "not suitable for task creation" when the chart itself is the problem: too trivial, answer labels already printed on it, illegible or cropped, or it requires outside data to solve.

Plan for roughly an hour on average, including revisions and resubmission. Some tasks take longer if you need extra iterations to stump the model.

Yes, as long as the prompt still has exactly one defensible answer (for example, "it never exceeds that value"). Two things to watch: don't word it so it sounds like you're assuming the event happens, and make the "none/never" outcome an explicitly acceptable answer.

Yes, as long as there's one defensible answer that's answerable from the chart alone, and the count isn't trivial or already labeled. Strong counting tasks usually involve real visual reasoning, like overlapping marks, ambiguous grouping, or legend/color categories, rather than just reading a printed number.

Reviewers grade against the Chartography Quality Rubric, across these dimensions:

  • Chart Quality
  • Prompt Realism
  • Reasoning Requirement
  • Self-Containment
  • Defensible and Verifiable Answer
  • Answer Correctness
C

Reviews and feedback

Bring it to your Pod Lead with the task link and the specific rubric dimension you think was mis-scored.

There's no standard appeal form. Take the same route as above: bring it to your Pod Lead with the task link and the dimension in question.

Everything for the Video Caption project: your Slack channels, what the project is about, and how to complete each section of a captioning task.

Channels

Project FAQ

Training and calibration together are capped at 5 hours.

Average handle time is 1.5 to 2 hours, depending on the complexity of the task.

There are no location restrictions. You can work from anywhere.

Everything for the Remedy project: your Slack channels, what the project is about, and how to complete a task.

Channels

About the project

Project overview coming soon

This section will cover what the Remedy project is, what a task looks like, and what good work looks like. Send the project overview and training document to your Workada contact and it will be added here.

Project FAQ

Can't find what you need? Post in #remedy-discussions, we're online most of the day.

A

Getting started

You'll need a login on workada.com (use the email tied to your interview) and access to GPT 5.6 Sol on High Reasoning. If you don't already have a GPT Plus subscription, ask your project lead. Workada covers a month of it and arranges it directly with you.

From your dashboard, open Tasks in the left sidebar, then claim "Create Remedy Task" under "Available to Claim." That opens the intake form where you'll submit your scrubbed input files, the source paper link, and your prompt.

The Project Remedy Instructions doc covers the workflow end to end. Read it in full before your onboarding session, and revisit it alongside the Prompt Writing and Golden Response guidelines as you go.

B

Writing prompts

The best prompts integrate multiple pieces of experimental evidence the way a domain expert actually reasons through a paper, not a single lookup. Two patterns reviewers consistently point to:

  • Ask the model to combine results across several figures, assays, or data files to reach one conclusion, rather than reading a single value off one source.
  • Give the model a real decision point, one where a plausible but wrong analytical route exists, and specify a fallback like NOT_IDENTIFIABLE for when the data doesn't support a clear answer.

Once a task has a well-formed prompt, the rest of the task tends to follow. Full guidance and examples live in the Prompt Writing Guidelines.

Specific enough that there's exactly one defensible way to do the analysis. Reviewers will flag a prompt if there are multiple reasonable interpretations. For example, several valid ways to define a fold-change calculation, which time points to require, or whether to rank by FDR or raw p-value all count as ambiguity. If more than one defensible answer exists, tighten the prompt until only one does.

That's still useful data. Log it in the Prompts With No Model Failure sheet:

  • Add a new row with the prompt, the model's response, and a shareable chat link.
  • You can leave the notes column blank for now.
  • Copy the task ID from the three-dot menu on the task and paste it into column A.

If the model seems close to failing, keep the same task open and try a revised prompt. If it's not close, it's usually faster to start a new task.

C

Scrubbing & data prep

Models are surprisingly resourceful about extracting an answer from metadata rather than the analysis itself. One task looked like a clean model failure until the reviewer noticed the file name itself contained the answer. Once it was renamed to something opaque, the model actually started failing as intended. Treat file names, sample labels, and folder structure as part of the prompt.

Yes. The task form has fields for both, plus a short text field explaining what you scrubbed. Reviewers use the original-to-scrubbed comparison to spot lingering answer-revealing details and to build better guidance on what to scrub next time.

Renaming files and sample identifiers has resolved this for other contributors, and it's often enough on its own to stop the model from recognizing the paper. If it persists, share the files you're using with your project lead in #remedy-discussions and they'll help identify what's leaking.

D

Working with the model

Yes, you can do this in the same conversation you'll submit. Just make sure the prompt and response you paste into the task are the ones from the first turn, not the debugging exchange that follows.

No. With memory off, a new task starts clean even if the earlier task hasn't been deleted. This is the recommended way to work through multiple prompts drawn from the same publication without cross-contamination.

Paste the visible response and include the shareable chat link. Hidden portions, like reasoning traces or "thought for X seconds" content, don't need to be pasted in for now; the chat link covers that.

E

Reviews & feedback

The [Remedy] Attempted Tasks and Reviews Tracking sheet is the source of truth for every task's status and score. It shows whether a task is good to go, needs changes, or needs a new sub-decision, and reviewer notes on what to fix live in column J.

  • Read the feedback in column J of the tracking sheet.
  • Questions about it go in #remedy-discussions, where each task gets its own thread.
  • Make the update in your Workada Dashboard, and confirm the model still fails after the change.
  • Once resubmitted, mark column M, "Updated By Attempter."

Reviewers do another pass from there, and this can take more than one round for a task to land.

The Golden Response Guidelines are a companion doc to the Prompt Writing Guidelines, clarifying exactly what belongs in the golden response field, added after reviewers noticed contributors interpreting that field differently. Worth reading in full before you submit your next task.

It means there's more than one defensible way to reach an answer, so the task doesn't have a single objectively correct result. The fix is usually to pin down the exact analytical decisions in the prompt itself: which comparisons to make, how to define the metric, which statistical threshold to use, and so on. See "Writing prompts" above.

F

Pay & referrals

Payouts run weekly. Your GPT Plus reimbursement is added to the payout following your subscription, not the same one, so expect it to show up a few days after you first mention the expense to your project lead.

Yes, referrals are welcome. You'll receive $150 per referral for anyone who completes an accepted task. Reach out to your project lead directly to make the introduction.

G

Support

Log it in the [Remedy Sites] Being Blocked by Workada Timer sheet and message your project lead directly. They'll get the timer updated as soon as possible.

Yes. Drop-in Zoom office hours run periodically, no need to stay for the whole session. Bring your task and your questions, get aligned with a reviewer, and head out once you're unblocked. Announcements for these go out in #remedy-announcements.

Coming soon

Resources for the Multimodal project are on the way. Your Pod Lead will let you know as soon as this page is ready.

Coming soon

Resources for the Lotus project are on the way. Your Pod Lead will let you know as soon as this page is ready.

Coming soon

Resources for the Multi-hop Reasoning project are on the way. Your Pod Lead will let you know as soon as this page is ready.

Coming soon

Resources for the Finance Acquisition project are on the way. Your Pod Lead will let you know as soon as this page is ready.

Coming soon

Resources for the Sheets Acquisition project are on the way. Your Pod Lead will let you know as soon as this page is ready.

Coming soon

Resources for the Sheets Artifact Collection project are on the way. Your Pod Lead will let you know as soon as this page is ready.

These are the standards for using Workada’s Dashboard and Slack. In short: use the platform only for legitimate task work, keep one secure account, submit accurate work, keep everything you access confidential, and treat everyone with respect. Read the full Code of Conduct →

On the Dashboard

In short: use the platform only for legitimate task work. One account, keep your credentials secure, submit accurate work, treat all task and customer data as confidential, don’t reverse-engineer anything, and don’t try to game the system. Everything you create belongs to Workada.
Account integrity
  • Keep only one account. Don’t create another after a suspension or termination without written permission.
  • Keep your login credentials secure, and never share, sell, or transfer access to your account.
  • Make sure your personal and payment information is accurate, current, and in your own name.
  • Complete any identity-verification request from Workada right away.
Performing tasks
  • Take on only tasks you can complete accurately and on time.
  • Follow the scope, specifications, and instructions provided with each task.
  • Don’t submit work that’s fabricated, copied from unauthorized sources, or done by someone else in your place.
  • Don’t try to game quality-control or verification. Repeated inaccurate work can lead to removal from projects or account deactivation.
Confidentiality and data
  • Treat all Workada materials, customer data, and task content as confidential.
  • Don’t download, copy, or store confidential information beyond what a task requires.
  • Don’t share task details or customer data with third parties, on social media, or anywhere else.
  • When a task ends or your account closes, delete all confidential information from your devices and cloud storage.
Intellectual property
  • Everything you create through the Dashboard belongs to Workada. Don’t keep copies or try to use, sell, or license it on your own.
  • Don’t submit content that infringes anyone else’s intellectual property.
  • Don’t use Workada’s materials for anything outside your tasks.
Software and systems
  • Use only approved browsers and apps. Don’t reverse-engineer, decompile, or extract source code from any Workada software or extension.
  • Don’t probe, scan, or test the security of the Dashboard or any connected system.
  • Don’t introduce malware, scripts, or bots that manipulate the platform or task results.
  • Report any security issue or suspected breach to support@workada.co right away.
Compliance
  • Use the Dashboard only for lawful purposes, following all applicable laws.
  • Don’t work from, or on behalf of, any country or party under applicable sanctions or export controls.
  • Don’t upload or submit content that’s illegal, discriminatory, harassing, or otherwise harmful.

Workada can investigate suspected breaches, hold payment during an investigation, remove you from projects, or suspend or deactivate your account, at its reasonable discretion under the Terms of Use.

On Slack

In short: be professional and respectful. Keep conversations work-related and on-topic. Never share task details, customer data, or confidential info in any channel. No spam, harassment, or solicitation. Workada can monitor and remove content, and serious violations can deactivate your account.
Professional standards
  • Communicate respectfully at all times. Personal attacks, insults, and hostile behavior aren’t allowed.
  • No harassment, discrimination, or bullying on the basis of race, gender, age, religion, disability, national origin, sexual orientation, or any other protected characteristic.
  • Handle disagreements constructively, and escalate anything unresolved to support@workada.co rather than letting it become a public argument.
  • Keep your language inclusive and appropriate for a professional, international audience.
Permitted use
  • Use Slack to discuss Workada tasks, workflows, platform updates, and professional development.
  • Post in the right channel, and keep messages concise. Off-topic or repetitive posts may be removed.
  • Don’t use Slack to solicit business, advertise third-party services, or recruit Contributors away from Workada.
Confidentiality and content
  • Don’t share confidential information, task details, customer data, materials, or work product, in any channel, DM, or thread.
  • Don’t post screenshots, documents, or files from the Dashboard unless Workada authorizes it.
  • Don’t post anything illegal, defamatory, obscene, threatening, or harmful, and don’t post spam or phishing links.
  • Don’t share or endorse misinformation about Workada, its customers, or other Contributors.
Privacy and moderation
  • Don’t share copyrighted material or third-party personal data without authorization.
  • Don’t record, screenshot, or distribute private conversations without everyone’s consent.
  • Moderators may monitor channels and remove content that breaks these standards. Report violations to support@workada.co.
  • Repeated or serious violations can lead to removal from the workspace, suspension, or account deactivation.

Running into a snag mid-task is frustrating: we get it, and we’re always working to make Workada better. Here’s how our support process works, plus a few tips that help us take care of you faster.

1

How to submit a support ticket

  1. You run into a problem while tasking, or you have a pay or timer issue.
  2. Open the Support module in the Workada Dashboard and email support@op.workada.co, or email that address directly using the email linked to your Workada dashboard.
  3. In the message, include your name, your email address, the Task ID (if needed), and a brief description of your issue.
  4. You’ll get an automatic reply confirming that a member of the support team will get back to you within the next 2 business days. Want to add more detail on something you’ve already submitted? Reply directly on the original ticket, don’t open a new one.
  5. Workada Support follows up on the same email thread to resolve the issue: either guiding you through a fix or resolving it directly on Workada’s end.
  6. No response after 5 days? Reach out to your Pod Lead and give them the ticket number, they can escalate your unresolved ticket.
Note If you’re running into the same issue repeatedly, message your Pod Lead directly instead of filing another ticket. A recurring problem usually needs a different kind of fix than a one-off ticket does.
2

Do’s and Don’ts at a Glance

DoKeep all your updates on the same email thread as your original ticket: it helps us help you faster.
Don’tNo need to open a separate ticket for the same issue; replying on the existing thread works better for everyone.
DoInclude the task ID and any references that can help us track down the problem when reporting a task issue.
Don’tTry not to report a task issue without an ID or supporting references: it makes it harder for us to locate the problem.
DoGive our support team the standard 2 business days to respond. We’re on it.
Don’tThere’s no need to assume no reply means no one’s working on it: we’ve got your message.
DoFeel free to loop in your Pod Lead if something feels urgent; they can flag it to a support agent directly on Slack.
Don’tIf you or your Pod Lead already reported the same issue, there’s no need to submit it again: one ticket is enough, and we’re already on it.
A quick note Our support agents do everything they can to get back to you as quickly as possible: often well before the 2-business-day window is up. Thanks for your patience, and for following the tips above; they really do help us take care of you faster. And thank you for everything you do as a Workada Contributor: we’re here for you.

From your application to your first task: here’s the full path, stage by stage, so you always know where you are and what comes next.

Team & Community

Your Slack home base. Join these channels, introduce yourself, and check in regularly.

Your first point of contact: your Pod Lead.
A

Sign Up

  1. Apply. Submit your application to Workada.
  2. Resume screen. Our team reviews your application and resume.
  3. Interview invitation. If you pass the screen, you’ll receive an email inviting you to an interview.
  4. Interview. You’ll meet with a member of our team over Zoom.
B

Onboarding

  1. Get access to Slack.
  2. Get your Pod assignment (more on this below).
  3. Complete Persona verification.
  4. Set up your bank account: allow extra time for this step; it’s the one that most often takes a few tries.
  5. Set up your timer.
C

Pod Assignment

Every Contributor is assigned to a Pod Lead: your go-to person for questions, guidance, and escalations.

  1. Slack channel. You’re added to your Pod Lead’s Slack channel.
  2. Welcome email. Your Pod Lead sends you a welcome email to get you started.
D

Training

  1. Finish training.
    If you passYou get access to calibration tasks.
    If you don’t pass yetReach out to your Pod Lead: it depends on the situation, and in some cases additional resources and training are available to help you get there.
  2. Calibration tasks. Complete your calibration tasks to fine-tune your ratings before real work begins.
    If you passYou’ll be able to continue with your project-specific tasks.
    If you don’t passReach out to your Pod Lead: it depends on the project, in some projects additional resources and training are available to help you get there.

    Not passing training and calibration may lead to offboarding.

E

Project Assignment

  1. Start tasking. You’re fully set up: browse available tasks and jump in whenever you’re ready. If you need help at any point, reach out to your Pod Lead.
Room to growThere’s a path forward within Workada: Contributors who deliver consistently and have great metrics can grow into Pod Leads.
OffboardingOffboarding can happen at any stage of the journey: either at your request (voluntary) or initiated by Workada (involuntary).