Getting started and setup
Complete the following steps to get fully set up:
-
Set up the Workada Timer
- Click "Set up" on the "Setup Workada Timer" card on your dashboard.
- Click "Download app" and install the timer on your computer.
- Open the app and click "Connect app" to link it to your account.
- When you see "You're all set," click "Continue."
Workada Timer app Overlay mini-timer -
Set up payments
- Go to the Payments tab on the left side of your dashboard.
- Follow the prompts to connect your Stripe account.
- Make sure this is done before your first payment cycle.
-
Accept Terms of Use
- Read and accept Workada's Terms of Use.
-
Complete identity verification
This takes about 2–5 minutes. Here's what you'll need to do:
- Upload a valid, non-expired government-issued photo ID: driver's license, passport, or state/national ID. Have both the front and back ready to photograph.
- Confirm your basic personal information (full name, address, date of birth): this is often extracted automatically from your ID.
- Complete a selfie/liveness check: you'll take a live selfie and may be asked to turn your head or look at different points on the screen. This confirms you're physically present and that your face matches your ID.
Most verifications are approved within minutes. Some cases may require manual review and take a little longer.
-
Complete training
Training is self-paced and completed on the Workada dashboard before you can access any live tasks. It's made up of multiple sections: each one builds on the last, with practice rounds followed by a short knowledge check. You need to pass each check to move forward.
Here's what you'll cover:
- What the work involves and what a good submission looks like
- How to spot AI artifacts: the subtle mistakes models make in hands, faces, text, and backgrounds
- The rating scale and how to use it
- How to write justifications that are specific, grounded, and consistent with your rating
- Each of the dimensions you'll judge: Instruction Following, Visual Quality, AI Appearance, Correctness, and the two Preference ratings
- Best practices for staying accurate and consistent across a task
The final sections put everything together with hands-on practice: spotting image oddities, evaluating how well an image follows a prompt, and comparing two images end to end. Passing the last knowledge check unlocks your tasking.
You’ll receive an email to your personal inbox with your Slack invitation. If you don’t see it, check your spam folder. Log in via Slack.com or the app.
You’ll be assigned to a pod channel, which is your primary space for day-to-day communication and support.
Once you receive your Slack invitation at your Workada email, log in via Slack.com or the app. You will be assigned to a pod channel: this is your primary space for day-to-day communication and support.
Pod Lead assignment happens automatically after you activate Slack. If you aren’t assigned to a pod within the next business day, email support@op.workada.co.
It's used to sign into Slack: there's no standalone inbox. All mail is automatically forwarded to your personal email, so please check your personal inbox regularly (including your spam folder) for any communications from Workada or your Pod Lead.
This is a known issue: some antivirus configurations interfere with the timer.
Use Google Chrome. Don’t use Safari or Firefox, they don’t reliably track your timer activity and can log active work as idle.
iOS, Android, and tablets, Chrome OS and Chromebooks, and Linux aren’t supported. An installed antivirus or the Firefox browser can also interfere with the timer. On a MacBook, use only Google Chrome.
Check the Project Resources tab, which covers most project-specific answers. If you can’t find what you need there, ask your Pod Lead.
Payment
As pay may vary by project, your rate is reflected on your dashboard under the Payments tab. You can find it at any time in the Payments tab on your dashboard.
Your payout is calculated as: (Eligible Hours × Hourly Rate) + Incentives (if earned)
"Eligible Hours" is the time that counts toward payment under our Pay Policy: this excludes idle time, activity on non-approved sites, and training time beyond your project's cap. See the Pay Policy and Timer tab for the full breakdown of what's eligible.
You can verify your expected payout at any time by checking your session history in the Sessions tab on your dashboard before raising a support ticket.
Payments are processed three times a week. The day your work is processed depends on when you complete it:
| Finish your work by | Processes on |
|---|---|
| Sunday, 11:59 PM PT | Monday |
| Tuesday, 11:59 PM PT | Wednesday |
| Thursday, 11:59 PM PT | Friday |
Once processed, Stripe typically takes 1–3 business days to deposit the funds into your account. If a processing day falls on a public holiday, payment will be moved to the last business day before it.
Yes. Workada never sees or stores your payment information. Your banking details are held directly and securely by Stripe, one of the world's largest payment processors.
Check your session history in the Sessions tab first to verify your logged hours. If something still looks off, email support@op.workada.co with the details.
Contributors are classified as independent contractors and are responsible for managing their own tax obligations. Workada does not withhold taxes on your behalf. Tax documentation is generated automatically through Stripe.
After you refer someone through your dashboard, your referral incentive shows up in your Payments tab within about two pay cycles. The clock starts once the person you referred completes their off-platform task.
Time tracking
All work must be tracked using the Workada Timer. You are responsible for clocking in before you start working and clocking out when you stop.
The timer also manages itself in a few ways: it automatically pauses after about five minutes without mouse or keyboard input, and it pauses whenever you're on a site or app that isn't approved for your project: you won't be able to resume until you switch back. Supporting tools like Slack and Google Search are fully paid as long as they're on the approved list, and training/calibration time is paid up to a cap that varies by project. For the full breakdown of what counts as paid time, see the Pay Policy and Timer tab.
You can see all of your tracked time in the Activity tab of your Workada dashboard, broken down by tasking versus other time, by day, and by the sites and apps you used.
Training and calibration are paid up to a cap specific to the project you have applied to. Once you pass that cap, the timer shows a warning and any time clocked in after that isn’t paid. If you run into this, reach out to your Pod Lead.
- Click Start before you begin any work or training
- Stop the timer when you finish your session
- Add a manual session if you forgot to start the timer
- Request a session edit if the timer recorded the wrong time
- Start working without the timer running: time not tracked cannot be paid
- Close the timer window thinking it stopped: it continues running in the background until you click Stop
- Run the timer when you're not actively working
- Forget to start the timer for training: it's paid time too
- Submit duplicate manual sessions for the same time period
Start your timer any time you are doing work-related activities: this includes training, tasking, attending meetings with your Pod Lead, or anything else related to your work on the platform.
Supporting tools like Slack, Google Search, and meetings are fully paid as long as they're on the approved tools list: see the Pay Policy and Timer tab for the full list. Note that AI assistants are not an approved tool for producing task output; using one won't count toward paid time and may affect your standing on the project.
Go to Dashboard > Sessions > + Add Manual Session. Enter the time window you worked and a brief description of the work completed. Submit it for review.
Manual sessions that are approved by Workada are paid out in the next scheduled payment cycle.
Yes: you can connect the Workada Timer on multiple computers if you work from more than one. To download it, go to the Sessions page or the Settings page on your dashboard and look for the download link. Note that you cannot run the timer on two computers at the same time.
Go to Dashboard > Sessions, find the incorrect session, and click "Request Edit." Adjust the time to what's correct, enter the reason, and submit. You can also contact your Pod Lead or email support@op.workada.co with the session details.
Check your Sessions tab at workada.com/sessions: your time is often saved on the backend even if it disappears from the app view. If the session is missing, add it manually via + Add Manual Session and contact support@op.workada.co with the details.
Go to Dashboard > Sessions, find the affected session, and click "Request Edit." Adjust the time to what's correct, enter the reason, and submit. If the session is missing entirely, add it manually via + Add Manual Session. If the issue persists, email support@op.workada.co with the details.
Neither. There's no minimum or maximum: you work at your own pace. What we care about most is the quality of the work you submit, not the number of hours you put in.
That said, when we review metrics, we do look at whether the time spent on a task reflects the quality of the work. Rushing through tasks is something we're able to identify, so take the time you need to do the work well.
We do look for Contributors who keep tasking regularly. Inactivity means going quiet on actual work, no submitted tasks for a stretch, not a quiet week in Slack or a slow reply to a message. If you need to step away from tasking for 7 or more days, let your Pod Lead know in advance.
Take the time you need: just let your Pod Lead know ahead of time so they can plan around your absence.
Use Stop when you're done working for the session; it ends your session and records the time. Use Pause for a short break you'll come back from; it keeps your session open and removes the break from your paid hours, so you don't have to stop and restart.
Both are fine to use. Pause just saves you from juggling separate start and stop periods for short breaks. Note that your balance won't show up in your Sessions history until you click Stop, and that paused time isn’t eligible for paid time.
This is a known issue: a few Contributors have had the chime keep ringing even after selecting "Still working." We're looking into a fix. In the meantime, opening the main timer window stops it.
We're aware that paused breaks can appear labeled as manual adjustments on your activity chart. We're working to ensure they're correctly labeled as ineligible time instead. This is a display issue only; your paused time is still correctly excluded from paid hours.
If you're on a MacBook, please use Google Chrome for tasking. We've confirmed that Safari doesn't reliably track activity on Mac; it can show active work as "Other" or "no activity detected," even when you're actively tasking. Switching to Chrome will track your time correctly and consistently.
Offboarding
Offboarding means losing access to continue contributing at Workada. There are five reasons you may be removed from the platform:
- Low quality work: if your quality falls below an acceptable threshold. Coaching always comes first, so reach out to your Pod Lead to run shadow sessions and improve your quality before it gets to this point
- Dishonesty: using AI to do your work for you, copying/pasting someone else's or your own prior work without doing the task yourself, spam, or time milking. First incident results in a warning; second results in offboarding
- Inactivity: going 7+ days without tasking (no submitted work) and without giving your Pod Lead a heads-up. This is measured by your tasking, not by your Slack activity or whether you reply to every message, and planned time off you've flagged in advance doesn't count.
- Community guidelines violation: egregious violations result in immediate offboarding with no warning
- Failing training or calibration: failure to pass either within the required attempts may result in offboarding
In all cases, you will be paid for all work completed up to your offboarding date, processed in the next scheduled payment cycle. Your Dashboard will remain available so you can check your Sessions and Payments history.
Notify your Pod Lead directly that you'd like to leave the platform. They will flag your departure to Workada Operations, who will revoke your Slack and Workada email access. Your Dashboard will remain available so you can check your Sessions and Payments history. You will be paid for all work completed up to your offboarding date in the next scheduled payment cycle.
Still have questions?
Email support@op.workada.co with your full name, a clear description of your issue, and any relevant details (task URLs, session times, screenshots, etc.) so the team can assist you efficiently.
At Workada, everything we build, every project, every partnership, every platform decision, comes back to three things. These aren't values we put on a wall. They're the principles that shape how we work, what we reward, and where we're headed together.
Quality
Quality is our north star. We don't trade great work for speed or scale. Every annotation matters because it shapes real AI outcomes, and because it determines what Contributors can access, earn, and achieve on the platform.
Quality isn't a standard we impose. It's the key that opens every door. The more you bring it, the more Workada brings back: more projects, more opportunity, more room to grow.
Community
Workada isn't just a platform, it's a community. We look out for our Contributors, and they raise the quality of everything we build. That's not a tagline; it's how this actually works.
When Contributors grow, in skill, in earnings, in impact, Workada grows too. We invest in the people who show up, do the work, and push each other to be better. This is a community that improves itself, together, and there's always room for people who want to be part of that.
Opportunity
At Workada, there's no ceiling on what you can achieve. The opportunity to grow, as an individual, in your career, and within the company, is real and it's ongoing.
High-quality Contributors always have projects waiting. Downtime shouldn't be something you worry about: that's our job. The pipeline stays full for those who keep the bar high. Show up, deliver, and the doors keep opening. That's the Workada promise.
Workada Time Tracking & Pay Policy Update
This policy explains how working time is tracked and paid. The principle is straightforward: you're paid for time spent actively doing approved tasking work. Below is how time is categorized and what counts toward payment.
Timer walkthrough
Here’s a short walkthrough of the timer and how your paid time is tracked.
What Counts as Paid
Project Tasking
Payment covers project tasks performed within the agreed and approved scope of your project. This work should take place on the approved tasking sites and apps.
If you are unsure whether an activity or tool is approved, consult your project instructions or contact your Pod Lead.
Training & Calibration
Onboarding time is paid up to a cap, specific to the project you have applied to. Workada's timer will show warning states as you approach the limit. Time spent on training and calibration beyond the cap is unpaid.
Supporting Tools
Certain tools support the work without constituting the work itself. Time spent on these tools is paid. The approved supporting tools are:
- Slack
- Zoom
- Google Meet
- Google Search
- Word
- Google Docs
- Image hosting sites used by tasks (for opening a task's images in a new tab to zoom or compare)
Use of AI Tools
Task output must be your own work. You may not use AI assistants (for example ChatGPT, Claude, Gemini, Copilot, or similar) to generate, complete, or fill in any part of your task deliverables. AI use will not be counted toward paid time.
If a specific project permits AI use for a defined purpose, that will be stated explicitly in your project instructions. Absent that, assume AI tooling is not allowed for producing work. Violations are treated as a quality and integrity issue, not a pay issue, and may affect your standing on the project.
Idle Time
To ensure that paid time reflects active work, the timer pauses automatically after five minutes without mouse or keyboard input, on any screen, application, or tab. When the timer pauses, you will be asked to confirm that you are still working. Acknowledging the prompt resumes tracking. Paused time is not paid.
Ineligible Activity
All other activity falls outside paid work, such as shopping, food ordering, streaming, and general personal browsing. The timer pauses automatically whenever you are not on an approved site, and that time is not paid. You will not be allowed to resume the timer until you switch to an approved tasking site or app.
Seeing Your Eligible Time
Your Workada timer will now separate tracked time into two categories, so you can see as you work what will and will not count toward payment:
- Eligible time: approved core tasking, supporting-tool time, and training time within your cap. This is the time that counts toward payment, subject to the standard review applied to all submitted work.
- Not eligible: idle time, activity on non-approved sites, and training time beyond your cap.
This breakdown is visible in your timer at any point during a session, so there are no surprises at the end of a pay period. If a block of time is marked not eligible and you believe that is an error, contact your Pod Lead.
Supported Browsers
We recommend using Google Chrome when tasking on approved project sites. If you're on a MacBook, Chrome is required: we've confirmed that Safari doesn't reliably track activity on Mac, and can show active work as "Other" or "no activity detected" even when you're actively tasking. Use of other browsers may affect the accuracy of your timer session.
Where you can start the timer
Your timer can only start from an approved work location. If you’re in a restricted area, the timer won’t start and you’ll see the message below. Restricted locations are set per project, so what counts as approved depends on the project you’re on.
VPN use isn’t allowed. A VPN can make your location look different from where you actually are, which can block your timer or flag your account.
Timer Best Practices
The Workada Timer is how your work turns into pay, so it's worth using well. Here's what counts as active work, what doesn't, and what to do if something looks off.
The Timer, and What Counts
Every Contributor must download the Workada Timer, which you'll find on the Workada Dashboard. Run it whenever you're actively working, and active work covers more than tasking:
Stop vs. Pause
Use Stop when you're done working for the session; it ends your session and records the time. Use Pause for a short break you'll come back from; it keeps your session open and removes the break from your paid hours, so you don't have to stop and restart.
Both are fine to use. Pause just saves you from juggling separate start and stop periods for short breaks. Note that your balance won't show up in your Sessions history until you click Stop, and that paused time isn’t eligible for paid time.
What the Timer Tracks
- Records which application or website is in focus, not what's on your screen
- Doesn't take screenshots, read window titles, or capture what you type
- Tracks mouse/keyboard input timing only, never content
- Records the country of your IP at session start
Task-Level Detail
- On Workada and your project’s webpage, the full page address is recorded and linked to the task you opened
- On every other site, only the site name is recorded (e.g. "wikipedia.org"), never the specific page
Overlay Colors
- Green means the focused window counts as eligible time
- Gray means it doesn't
- Eligible time includes approved tasking sites, supporting tools (Slack, Zoom, Google Meet, Google Search, Word, Google Docs, Email), and training time up to your cap
Switching Windows Briefly
- Quick switches won't cost you time thanks to a short grace period before auto-pause
- Extended time on non-approved sites pauses the timer and isn't paid
- Think a block was marked wrong? Contact your Pod Lead
Automatic Pausing
- Pauses on its own after five minutes with no mouse or keyboard input
- You'll be asked to confirm you're still working; acknowledging it resumes tracking
Do's and Don'ts
Do
- ✓Run the timer only while actively working: anything in the list above counts.
- ✓Pause the timer the moment you step away for anything personal, even a few minutes.
Don't
- ×Leave the timer running during personal browsing, socializing, or idle waiting.
- ×Treat Slack as unlimited paid time. Only work-related Slack belongs on the timer: chat and socialize on Slack once the timer is off.
Additional Guidelines
- You can check your hours yourself on the Dashboard's Sessions tab. Pay is hours worked × your rate.
- Forgot to start your timer while working? You can add a manual session in your timesheet.
- Left the timer running by accident while not working? You can edit your session to remove the time.
- If a warning email lands, there's no need to panic or stop working. Make sure you're using the timer when you're supposed to: document your side, and escalate through your Pod Lead rather than stress.
- One correction request is enough: a single support email or Sessions-tab request does the job, and duplicates slow things down.
Known Issues
- Chime not stopping after you confirm: a few Contributors have had the idle-check chime keep ringing even after selecting "Still working." We're looking into it; opening the main timer window stops it in the meantime.
- Paused time showing as a "manual adjustment": paused breaks may currently be labeled as manual adjustments on your activity chart. We're working to ensure they're correctly labeled as ineligible time instead. This is a display issue only; your paused time is still excluded from paid hours.
0 of 3 reviewed
We've confirmed that Safari doesn't reliably track activity on Mac; it can show active work as "Other" or "no activity detected," even when you're actively tasking.
If you're on a MacBook, please switch to Google Chrome for your timer to track correctly and consistently.
Referrals just got easier! You can now submit and track referrals directly from your Workada Dashboard: no separate form needed.
Here’s how it works:
- Once you’ve completed 5 hours of work on the platform, the Referrals section will appear automatically in your Dashboard.
- Submit your referrals there and follow their progress in real time. All tracking now lives in the platform, so you’ll always know where things stand.
- Your referral gets an automatic email inviting them to sign up. Once they’re approved and complete their first project task, you both earn a reward.
Rewards at a glance:
- Refer a Contributor: you earn $50 and they earn $20.
- You can have 5 referrals at a time, and the limit refreshes every 30 days.
One more thing: we’re retiring the old referral form, so please use your Dashboard for all new referrals going forward. Any referrals you’ve already submitted through the form will still be tracked: nothing gets lost.
Thank you for helping grow this community. Happy referring!
You’ll find announcements from the Workada team on this tab: updates, news, and changes involving the Workada Contributor community. Try to make it a habit to check here before you start tasking. Once you’ve read an announcement, mark it as reviewed: new ones will arrive unmarked.
Everything for the Multimango and Video Eval project: your Slack channels, how to evaluate image tasks, the video training and calibration, and project-specific questions.
Channels
Onboarding for the Video Comparison and Evaluation project.
Multimango announcements from the team.
Connect with other Multimango Contributors.
Discuss Multimango tasks and share tips.
Company-wide Workada announcements. Check daily before you task.
Task won’t load, throwing an error, or Multimango seems down? Use the Report - Task issues and outages button. Anything else is rare, but if it doesn’t fit, use Report - Other. This is the fastest way to get a task issue fixed: reports go straight to the project team, who investigate and get you back to tasking as quickly as possible.
Screenshots help (just leave out anything personal or payment-related), and you’ll get a quick confirmation once it’s in, no need to follow up.
- Read the prompt carefully: including implied context like "for a school project" or "should look like an ad"?
- Listed what must be included: objects, people, text, layout, colors, style?
- Noted what shouldn't have changed from the prompt image: did anything change anyway?
- Zoomed into both images (hands, faces, text, edges, backgrounds), not just glanced at the whole?
- You're comparing A vs. B, not just describing one of them?
- Checked whether either output changed something the prompt never asked to change?
- Does it follow the prompt? Everything requested is there, in the requested style, layout, and text: and the rest is untouched?
- Does it make sense? Logical scene, realistic proportions, natural behavior, sensible context?
- Does it look polished? Could you recreate this on purpose in Canva or Photoshop: or is it distorted, messy, obviously AI?
- Is it correct? Objects, anatomy, logos, text all accurate (verified against reality where checkable), judging only what applies to this prompt?
- Anatomy: distorted or misshapen bodies, missing or extra limbs, merged objects, impossible proportions?
- Text: garbled, misspelled, or gibberish lettering anywhere in the frame?
- Repetition: cloned faces or duplicated objects hiding in crowds and backgrounds?
- Physics: unrealistic textures, lighting, reflections, or shadows?
- You compared both images and explained why one wins, with specific examples: no vague statements?
- You used the prompt's own wording and descriptive language: distorted, misshapen, cohesive, realistic, whereas?
- Your comments support your ratings: one consistent case from start to finish?
- Your Overall Preference matches the issues you described in your comments?
- Proofread, inside the character limit, timer running, session confirmed?
| Question | Ask yourself | Look for |
|---|---|---|
| Overall Preference | Which image is better overall? | Prompt fit, quality, and errors weighed together |
| Instruction Following | Did it do what the prompt asked? | Missing items, wrong style, wrong colors, extra clutter |
| Visual Quality | Is it well presented? | Blur, lighting, cropping, focus, composition |
| AI-Generated Appearance | Does it have AI-looking mistakes? | Extra fingers, warped faces, melted objects, gibberish text |
| Correctness | Are checkable details right? | Spelling, dates, labels, counts, maps, facts |
What to look for
- Which image best fulfills the user’s request overall
- Which strengths and weaknesses matter the most
- Whether one major issue outweighs several minor ones
- Overall visual appeal, clarity, and presentation
Keep in mind
- This is a holistic judgment made after your detailed review, weighing everything you found across all dimensions: not a quick first impression, unless the task instructs otherwise
- Not every category carries the same weight in every task
- It’s not a category count: give a well-supported answer rather than tallying which image “won” more questions
- You’ll rarely have a true tie
- Avoid ties unless the images are genuinely very similar overall
What to look for
- Does it match the requested style: realistic photo, illustration, edit, greeting card?
- Are the requested objects, people, attributes, and edits all there?
- Are any required details missing or incorrect?
- Were unnecessary or conflicting elements added?
- For editing tasks: were unchanged parts of the original preserved?
Keep in mind
- Start with the big picture: what was the user actually asking for?
- Not every requirement carries the same weight; think about which issues have the biggest impact
- Added details aren’t automatically wrong if they fit naturally and don’t conflict with the prompt
- Judge the user’s overall intent, not just a checklist
What to look for
- Sharpness and focus
- Lighting and contrast
- Composition and framing
- Resolution and texture quality
- Realistic perspective and depth of field
- Editing seams or blending issues, when applicable
Keep in mind
- Look at the whole image first, then zoom in: small issues can matter
- Blur isn’t automatically a flaw; motion blur or shallow depth of field may be intentional
- A technically sharp image isn’t always better if composition or presentation is weaker
- Don’t over-penalize small flaws that don’t significantly affect the image overall
What to look for
- Distorted anatomy, especially hands and fingers
- Warped or asymmetrical faces
- Gibberish or malformed text
- Duplicated elements and repetitive patterns
- Melted or merged shapes; objects that didn’t form properly
- Impossible geometry or perspective; unnatural textures
Keep in mind
- Focus on artifacts, not on which image looks more realistic
- One major artifact may matter more than several smaller flaws
- Stylization alone doesn’t mean an image is AI-generated: illustrations and cartoons can be fully coherent
- If neither image has noticeable artifacts, don’t hunt for issues that aren’t there; an equal rating may be appropriate
What to look for
- Incorrect dates, numbers, or measurements
- Wrong labels, signs, or logos
- Spelling errors
- Impossible layouts or spatial relationships
- Historical, geographical, or domain-specific inaccuracies
Keep in mind
- Use Google to verify when needed: correctness is about details that can be objectively checked
- Missing requested elements belong under Instruction Following
- Gibberish or unreadable text usually belongs under AI-Generated Appearance
- If there are few checkable details, this question may not be a major factor for that task
Some observations could reasonably fit more than one question. Think about the main issue you’re describing and place it where it fits best.
| If the issue is… | It usually belongs under… |
|---|---|
| Prompt asks for 5 people, image shows 4 | Instruction Following |
| Prompt asks for red, image shows blue | Instruction Following |
| Sign says “July 32, 2026” | Correctness |
| “Miami” is spelled “Maimi” | Correctness |
| Text is unreadable or random symbols | AI-Generated Appearance |
| A hand has seven fingers | AI-Generated Appearance |
| Image is blurry or pixelated | Visual Quality |
| Subject is cropped too tightly | Visual Quality |
| A person is missing from the scene entirely | Instruction Following |
| A logo looks almost right but isn’t | Correctness |
| Fabric and skin have blended together | AI-Generated Appearance |
| The same face appears 3 times in a crowd | AI-Generated Appearance |
| Shadows point in conflicting directions | AI-Generated Appearance / Visual Quality |
| A building is architecturally impossible | AI-Generated Appearance / Correctness |
- Keep the prompt in mind for every question. User intent is your context: did they ask for an illustration, a realistic image, a card, a studio shot? Bright sunlight or a rainy day? Warmth, sadness, shock?
- Zoom in: every time. You can’t evaluate these tasks properly at full-frame view. Small issues decide close calls.
- Google what you don’t know. If the prompt mentions something unfamiliar, look it up on official sites. Answering with verified background knowledge makes your data far more accurate and valuable.
- 250 characters go fast. You won’t fit every detail, so lead with the ones that best support your answer for that question.
- Keep your answers separate. The questions are mostly about different things, so different details should support each one: avoid repeating the same observation everywhere.
- Avoid circular reasoning. “It looks more realistic because it’s more realistic” isn’t a justification. Name the tells: the plasticky plants, the gibberish text on the book.
Instruction Following asks: did the model do what the prompt asked? Correctness asks: is what the model created factually accurate? An image can pass one and fail the other: here’s how the same scenario scores on both.
| Scenario | Instruction Following | Correctness |
|---|---|---|
| Prompt asks for a portrait of a man in a blue coat; the image shows a black coat. | Error The prompt specified blue and the model ignored it. |
OK The image itself is accurate: the model just didn’t follow the requested color. |
| Prompt asks for a map of the Caribbean with its islands labeled; the map appears, but some labels are wrong. | OK A labeled map of the Caribbean is what was asked for, and that’s what was generated. |
Error Some island labels aren’t factually accurate. |
| Prompt asks for a square Save-the-Date for Juan & Mitchell reading “Save the Date – October 13th”; the layout is right but it says “Michelle” and “Octobre.” | Error The prompt asked for specific names and dates, and the model failed that request. |
Error The name and month are misspelled. |
| Prompt asks for a chess set with red and white pieces; the image shows a correct chess set with black and white pieces. | Error The prompt specified red and white pieces and the model ignored the requirement. |
OK The chess set itself is correct. |
| Prompt asks for a bird’s-eye view of Times Square at night; the setting and view are right, but buildings are misshapen and signage has text errors. | OK The model created the correct setting and time of day. |
Error The malformed buildings and text errors need to be pointed out. |
A lot of AI images look convincing at first glance, but they almost always contain subtle mistakes: and your job is to find them. Don’t just ask whether an image looks good. Ask what’s wrong with it, and explain it clearly.
- ✓Sharp: the subject is clear and well-defined. “The image is sharp, and facial details are clearly visible.”
- ✓Crisp: fine details render cleanly with strong clarity. “The feathers appear crisp with excellent texture detail.”
- ●Soft: details are slightly smooth or less defined, not necessarily wrong. “The image is slightly soft around the edges of the subject.”
- ×Blurry: important details are hard to see due to lack of focus. “The image appears blurry, particularly around the face.”
- ×Out of focus: the focus is placed incorrectly. “The subject is out of focus while the background appears sharper.”
- ✓Properly exposed: highlights and shadows are balanced. “The image is properly exposed and retains detail in both bright and dark areas.”
- ×Overexposed: too bright, losing detail. “The sky is overexposed and lacks visible cloud detail.”
- ×Underexposed: too dark to distinguish details. “The subject’s clothing is underexposed and difficult to see.”
- ×Blown highlights: bright areas lose all texture. “The highlights on the white shirt are blown out.”
- ×Crushed shadows: dark areas lose visible detail. “The shadows beneath the table are crushed and contain no visible detail.”
- ✓Natural lighting: light appears realistic and believable. “The natural lighting creates a realistic outdoor appearance.”
- ✓Soft lighting: smooth transitions, minimal harsh shadows. “Soft lighting creates flattering skin tones.”
- ✓Well-lit: the subject is clearly visible. “The image is well-lit and easy to evaluate.”
- ×Harsh lighting: strong shadows and intense highlights. “Harsh lighting creates distracting shadows on the face.”
- ×Flat lighting: the image lacks contrast and depth. “Flat lighting causes the subject to blend into the background.”
- ×Poorly lit: insufficient or inconsistent lighting. “The scene is poorly lit, making details difficult to assess.”
- ✓Strong contrast: clear distinction between light and dark. “The strong contrast helps the subject stand out.”
- ✓Balanced contrast: light and dark areas are well controlled. “The image demonstrates balanced contrast throughout.”
- ×Low contrast: tones lack separation. “The image appears dull due to low contrast.”
- ×Muddy tones: tones blend together without definition. “The background contains muddy tones with little separation.”
- ✓Accurate color balance: colors appear realistic and natural. “Skin tones exhibit accurate color balance.”
- ✓Vibrant: colors appear rich and lively. “The vibrant colors enhance visual appeal.”
- ●Muted: colors appear intentionally subdued. “The muted palette creates a softer mood.”
- ×Oversaturated: colors are excessively intense. “The flowers appear oversaturated and unrealistic.”
- ×Undersaturated: colors are washed out. “The image looks undersaturated and lacks visual impact.”
- ×Color cast: an unwanted tint affects the image. “The image has a noticeable blue color cast.”
- ✓Rich texture: surface details are clearly visible. “The wood grain exhibits rich texture.”
- ✓Fine detail: small details survive a close look. “Fine detail is preserved in the hair.”
- ✓Clean rendering: objects appear smooth and defect-free. “The image demonstrates clean rendering throughout.”
- ×Artificial smoothing: details are unnaturally softened. “Artificial smoothing reduces skin texture.”
- ✓Good depth: a realistic sense of space. “The image demonstrates good depth between foreground and background.”
- ✓Strong subject separation: the subject stands out clearly. “Strong subject separation draws attention to the focal point.”
- ●Shallow depth of field: sharp subject, blurred background. “The shallow depth of field effectively isolates the subject.”
- ×Flat appearance: the image lacks dimensionality. “The scene appears flat due to limited depth cues.”
- ✓Well-composed: elements are arranged effectively. “The image is well-composed with a clear focal point.”
- ✓Balanced framing: the subject sits well within the frame. “The balanced framing keeps attention on the subject.”
- ✓Strong focal point: attention is directed clearly. “The strong focal point immediately draws the eye.”
- ×Awkward crop: important parts of the subject are cut off. “The awkward crop removes part of the subject’s arm.”
- ×Distracting composition: unnecessary elements compete for attention. “The bright object in the background creates a distracting composition.”
| Instead of… | Try… |
|---|---|
| “It looks weird.” | Low contrast |
| “Something feels off.” | Flat lighting |
| “Bad quality.” | Soft focus |
| “It doesn’t pop.” | Weak subject separation |
| “Too bright.” | Overexposed highlights |
| “Washed out.” | Undersaturated colors |
Do
- +Zoom in inch by inch: most issues hide at normal size.
- +Search when you need to verify a fact, flag, or count.
- +Echo the prompt's own wording in your answer.
- +Build a case: name specifics in both images.
- +Stay consistent: your Overall Preference should match the issues you name.
Don't
- ×Copy-paste responses across dimensions or tasks.
- ×Use AI to help you answer.
- ×Contradict yourself: a hand you flagged as mangled shouldn't win naturalness.
- ×Settle for a tie: make the call whenever the images let you.
- ×Write "A looks more polished" and stop there.
→ Open the full Task Guidelines document
- Text To Image Compare
- Reference-to-Image Ranking
- Search-Grounded Image Generation Eval
- realism-hc-elo
- Realism Quiz
- Cv2 R2i Heldout
- Text-to-Video Ranking
- Image-to-Video Ranking
- Omni R2v Elo
- Omni Tts Elo
- Text-to-Audio - Audio-Only Evaluation
- Multi-Capability Model Ranking
- Ad Creative Overall Ad Suitability (Q0)
- Ad Creative-Visual Flaws (Q2)
- Ad Creative – Photorealism (Q3)
- Ad Creative– Visual Appeal (Q4)
- Ad Creative – Text Flaws (Q5)
- Ad Creative – Text Style (Q6)
Overview
Everyone starts on image tasks. Once you qualify, you'll train on video and then get video comparison tasks mixed in with your image work. Here's how you become eligible, what training and calibration look like, and how to score each comparison.
How you qualify
You become eligible for video training once you’ve built a track record on image work: 15 active hours on multimango.com image tasks. Once you hit 15 image hours, you’re moved into video eval setup automatically.
What you’ll do
Qualifying is two steps:
- Video training walks you through the video task and the quality dimensions you’ll be judging. You’ll need at least 75% to move on to calibration, and you get two attempts on each knowledge check.
- Calibration: you compare two AI-generated video clips and judge them across the task’s 9 quality dimensions (video and audio), picking the better clip overall. You complete 4 calibration tasks, and 2 passing tasks qualifies you for video tasking.
Your time cap
You get 3 paid hours for training and calibration, and you complete your 4 calibration tasks within that time. Your timer warns you as you approach the cap, and time beyond 3 hours isn’t paid, so pace yourself. It’s built to fit inside the window.
While your Pod Lead reviews
You can’t continue tasking until you’ve completed all 4 calibration tasks, so get them done within your time cap. Once they’re in, you can go back to tasking on Multimango while you wait for your Pod Lead to grade them. Once you pass, video tasks unlock. To be clear, you don’t switch to video-only: you’ll see the same standard mix of image and video tasks you had on your dashboard before.
How to score: slightly, strongly, or tie
Each task shows you two AI-generated videos. For every dimension, decide which side wins and by how much — it’s a per-dimension call, made fresh on each line.
One video is better, but the gap is small. You can name the specific advantage — yet another careful rater could reasonably call it a tie or lean the other way.
One video is clearly better on that dimension and you’re convinced of it. Not a lean — a conviction you’d hold even if pushed back on.
The dimension genuinely doesn’t separate them: no relevant advantage, or strengths and weaknesses offset. The videos don’t need to be identical.
Slight vs. strong at a glance
| What to consider | Slightly prefer | Strongly prefer |
|---|---|---|
| Why it wins | Better, but by a fairly small margin. | A clear, immediately visible advantage separates the two. |
| Winner’s flaws? | Yes — and they may make the call closer. | Maybe, but not enough to change the outcome. |
| The other video? | Has problems, but balances them with strengths. | Its problems widen the gap. |
| Room to disagree? | Yes — a careful rater could call it a tie or flip it. | Probably not — a careful rater should land in the same place. |
One difference from image work: on video tasks, Overall Preference is your gut-feel call, unlike image tasks where it’s an in-depth final choice weighing all dimensions.
Common Calibration Errors
These are the most frequent issues identified in calibration and quality reviews:
- Same pick every time: choosing the same video for every dimension instead of judging each one on its own.
- Wrong dimension: writing about something that belongs under a different dimension, like describing visual quality under Instruction Following.
- Rushed or generic: a justification that could apply to any pair of videos, with no specific detail or moment cited.
- Contradiction: writing a rationale that argues for the video you didn't actually pick.
- No explanation or cut off: leaving an explanation box blank, or submitting a justification mid-thought.
Project FAQ
Multimango
You'll complete at least 10 calibration tasks, labeled "[Calibration check]: Compare two images," with up to 20 attempts total. The 70% is based on the share of tasks you pass, not an overall star average, so aim to pass at least 7 out of 10. This 70% bar is for image calibration; initial training requires 75%, and Text-to-Video (video) calibration uses a 2-of-4 bar instead. These numbers can change, so check with your Pod Lead if you're unsure.
If you're not at 70% after your first 10 attempts, you'll be prompted to complete additional training and reach out to your Pod Lead: you'll still have attempts left at that point, so that's a great time to ask for support. If your first 5–10 calibration tasks aren't reaching a 70% passing grade, reach out to your Pod Lead for support right away so you don't risk being offboarded from the project. If you don't reach 70% within all 20 attempts, you'll be offboarded from the project. This is the Failed calibration reason.
Only after your Workada dashboard tells you to. Do not sign up on multimango.com on your own: you will not be approved unless you go through the dashboard flow and click the "I created my account" button.
- In your Workada dashboard, find the "Getting Started" card and click Set up on "Create Multimango account."
- In the pop-up, click Open multimango.com. Keep the Workada window open: you'll come back to it.
- On multimango.com, enter your workada.co email (shown on your dashboard when you click Set Up) and click Continue. Note, this should be workada.co, NOT workada.com!
- Create a password (at least 8 characters) and click Continue.
- Check your email for a 6-digit verification code, enter it, and click Continue.
- You'll see a "Welcome!" screen. Return to the Workada window and click "I created my account."
- Wait for approval (done manually: typically within 72 hours).
- Once approved, select "Text To Image Compare" as your task type, and you're working.
Always use your workada.co email, not a personal address. Accounts created with a personal email can't be approved.
Your code is sent to your workada.co address, then forwarded to the personal email you originally signed up with: check both "All Mail" and "Spam" in that inbox. If it's expired or never arrives, use the "Resend" button to get a new one.
Approvals happen manually, typically within 72 hours. Once you're approved and your account is active, you will get an email, and you'll see a "Tasks" option appear in the Multimango sidebar. There's nothing else you need to do while you wait.
Work whichever of these four task types is showing on your Multimango dashboard:
- Text to Video
- Multi-Capability Model Ranking
- Text to Image Compare
- Reference to Image Ranking
These can appear and disappear at random, so you may not see all four at once: but everyone should have at least one available at any given time. If your usual type isn't showing (or won't let you pick up a task), work one of the others that has available tasks rather than sitting idle.
Prioritise video and image tasks. Other task types sometimes appear on your dashboard, such as UD Caption or Caption Quality Ranking. Only work those if no video or image tasks are available to you.
Task availability for a given category can fluctuate. If one or more tasks have the message "No Active ELO Evaluations" when you attempt to access them, go back and try another category. Remember to refresh your screen to make sure no new tasks for our priority categories appear. Do not email support about receiving the No ELO message on Text to Image Compare. Contact your Pod Lead for guidance if you consistently see it across all other task categories, especially multiple categories.
Tasks can time out, and the type may temporarily drop off your dashboard afterward. Refresh and check the other available task types: it typically reappears. We're aware tasks sometimes time out sooner than they should and it's on our fix list.
Quality is also read from behavioral signals, like speeding through tasks, taking too long, or picking the same side (for example, all A) many times in a row. You'll usually get a warning before you're blocked, so treat one as a cue to slow down and refocus on accuracy.
This message can have more than one cause. It’s shown by Multimango, not by Workada, and the same wording appears in several different situations:
- We moved you between Multimango annotator groups. When that happens, Multimango drops whatever task you had claimed at that moment and your submit fails. This has nothing to do with your work.
- Multimango paused your account based on its own behavioral checks. Multimango has described these to us as patterns like moving through tasks unusually fast, leaving a task open for a very long time, or selecting the same side (for example, all A) repeatedly. Some of these pauses lift on their own after roughly 24 hours. Others don’t.
- You’re in video calibration. Until your calibration tasks are graded, you can’t pick up image tasks, and this message is what you see when you try.
If you’ve seen this message, don’t assume it’s a verdict on your accuracy, and don’t assume you’re being offboarded. The pause lives entirely in Multimango’s system. It doesn’t appear anywhere in Workada, which means your Pod Lead and our Ops team genuinely can’t look up why a particular account was paused.
If you get the block message:
- Don’t retry the same task. It won’t go through.
- Refresh and claim a new task. For most people that clears it.
- Wait about 24 hours before assuming it’s permanent.
- If it follows you onto new tasks, or no tasks appear at all, message your Pod Lead with your email, the task type, and roughly when it started. Please go through your Pod Lead rather than the support email for this one, so we can tell an account-level pause apart from a group change.
Let your Pod Lead know. They can review it and escalate it for you if the QA looks incorrect.
The only locations you can work from are the US and Canada, plus US territories by default (PR, GU, VI, AS, MP).
Review the dimensions first, then split your answers up by category, calling out different inaccuracies for each dimension instead of repeating the same point.
Instruction Following asks “did the image do what the prompt asked?” You’re checking the prompt’s requirements: subjects, objects, and actions; counts; attributes like color and clothing; style and setting; and negative constraints like “no text” or “sun not visible.”
Correctness asks “are the checkable facts and structures right?”, whether or not the prompt asked for them: spelling and legible text, accurate counts, anatomy and physics (limbs, shadows, reflections), and factual labels, logos, and flags.
Rule of thumb: Instruction Following is prompt faithfulness; Correctness is verifiable errors.
Move on to the next task. Once a task is closed you can’t go back to the same prompt.
Write a sentence for every applicable dimension, on every task type. Some tasks show all the dimensions to judge individually with a single justification box at the bottom: even then, cover each dimension you scored rather than writing one general comment.
It also helps a QA reviewer understand why you preferred one response over the other.
Refresh the page to get a new task. Don’t post screenshots of the material in community channels: sharing it exposes other Contributors to the same content. We share reported examples with the Multimango team and are pushing for better filtering on their end, but unfortunately can’t make any guarantees.
No, these don’t need to be flagged to the MM Task Issues channel.
Tasks and workflow
Tasks are available through the Task tab on your Dashboard and/or on Multimango (see next section for more details). Once your onboarding session is complete, your training is passed, and your account is fully set up, you'll be able to browse and pick up available tasks.
| Status | What it means |
|---|---|
| Not started | The task has been assigned or made available, but work hasn't begun. |
| In progress | You've picked up the task and are actively working on it. |
| Submitted | You've completed and submitted the task for review. |
| Pending review | Your submission is queued with the quality team: this is normal. No action needed unless it's been over 48 hours. |
| Reviewed | The task has been reviewed. You may receive feedback or a score. |
There's no strict time limit, but aim for around 12–17 minutes on Text To Image Compare tasks and around 20 minutes on video tasks in the Multimango project. The 12–17 minute range applies to Text To Image Compare specifically: other task types, such as Reference to Image Ranking or the caption task types, routinely take longer and aren't held to it. Tasks on multimango.com are capped at 20 minutes, and image and video calibration tasks at 30 minutes, so going over these can affect your metrics. Other projects may differ. Trust your gut on the obvious calls, zoom in where it matters, and only fact-check claims that are actually checkable: a count, a label, a spelling. Don't go down rabbit holes.
Quality is what matters most, so take the time you need to do the work well. As you build experience, you'll naturally get faster without losing that quality.
Once you pass training, you'll need to complete calibration tasks first: see the Multimango section below for details.
Yes, either way works.
There isn't a rule that fits every task, so use your judgment. Slightly is the default, but don't hold back strongly when one option clearly wins on that dimension and you can defend it.
Refresh the page to load a new prompt.
Long prompts on image and video tasks can eat your whole handle time. Work them in this order:
- Review the outputs on their own, before you read the prompt.
- Read the prompt and identify the primary prompt: what the output basically has to be.
- Skim the rest, separating the details that matter from the minor ones.
- Make your selection on that basis.
- Write your justification against the details you judged important.
If a prompt is so long that this still isn't workable, let your Pod Lead know.
Take the time the task actually needs. Where it is unreasonable to complete a task properly inside the limit for that type, quality comes first.
You won't be blocked for going long while you're working in good faith. That only becomes a problem when the time is far above average and starts to look like time milking.
Yes, as long as your Workada timer still shows MultiMango as your active tab. Don't use one if it significantly increases your handle time for little or no gain in quality.
Offboarding
Yes, you may apply if you’d like. Keep in mind that the Data Labeling Specialist role has the easiest training and calibration, so other roles may be more challenging.
Video Eval
After you’ve completed your 4 calibration tasks on workada.com, you can resume tasking on multimango.com until your Pod Lead has reviewed them.
There’s no opt-out. Everyone starts on image tasks and moves into video once they qualify. Video eval is an explicit request from the Multimango team, and purposely failing training to stay on image tasks is grounds for offboarding.
Switch to the Video Eval project, open the Training tab, and start your timer.
No. You’ll see the standard mix of tasks on your dashboard, just like before: image and video. (A separate video-only tier exists for top performers over time, but training alone doesn’t switch it on.)
Complete 4 calibration tasks; 2 of them passing qualifies you for video tasking.
For now, you keep doing image tasks as normal. Failing video doesn’t remove your image work.
Gold Standard, an easy benchmark task that’s auto-graded to check your accuracy on image work. It’s one of the signals used to assess your quality, though video eligibility itself is based on your image hours.
We’re working on giving you that visibility directly. For now, no action needed: keep doing your image tasks, and once you hit 15 image hours you’re moved into video eval setup automatically.
Your timer warns you as you approach it. Time past the cap isn’t paid, so aim to finish your 4 calibration tasks within the window.
Your Pod Lead grades your calibration and shares guidance.
Only your Text-to-Image hours count toward the Image Active Hours used for video eligibility. Until you’ve passed video training, other task types don’t count toward that total.
No. When a task only includes Audio Quality and Sync, evaluate just those two, not audio instruction following.
Once you pass, you move into the general task pool and the client routes whatever task types it wants to you, that could be image, video, or another type. Before you pass, you're limited to video tasks only.
Aim to stay under the 20-minute cap, you'll be kicked off a video task past it. Text-to-video tasks usually have just 1 to 4 dimensions to score (never more than 6), so pacing yourself is manageable.
The two axes no longer overlap:
- MTQ (Motion & Temporal Quality) has been narrowed to traditional motion-quality issues only: flickering, smoothness, and camera jitter.
- LAIG (Less AI Generated) has been expanded to cover physics issues, phasing in and out of reality, extra limbs, and other generative artifacts.
Physics and AI-related artifacts used to get scored under MTQ. They now belong under LAIG. The change separates the two axes and lines annotators up with auditors.
Training, calibration answers, and the tearsheet are being updated to match. Until then, use the training and calibration examples as your reference.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
Everything for the Sheets project: your Slack channels and the questions we hear most.
Channels
Sheets FAQ
Getting started and setup
Welcome aboard. Here’s your path from here:
- Verify your identity.
- Set up Slack.
- Set up your Workada Timer.
- Accept the terms and conditions.
- Complete the training modules.
- Complete the calibration task.
- Complete your first two tasks.
Your first two tasks go to a reviewer to check and accept. Once they pass, you’re cleared to claim live tasks from your Dashboard.
Most Contributors finish in under 2 hours, and we compensate up to 2 hours of training time. If you need a little longer to feel ready, that’s completely fine. Anything past the 2-hour mark just isn’t compensated.
Almost there. Your first two tasks go to a reviewer first. Once they’re accepted, you’re cleared to claim live tasks from the main pool on your Dashboard.
Yes.
Answer coming soon.
Tasks and reviews
Hang tight, this is normal. Your first tasks are assigned to a reviewer by hand, so they can take longer than pool tasks. If yours has been waiting more than 5 business days, reach out to your Pod Lead and they’ll help move it along.
We’re actively working to improve these metrics. For now, here’s what we capture for each task:
- Claude score on first attempt. The review rating your task gets the first time you submit it, as it moves from edit_task into task_review. The review system scans the prompt and workbooks for factual contradictions and critical coverage gaps, which makes it the cleanest read on how close your task is on the first try. A score of 3 or below sends the task back for a redo; a 4 or higher, with the rubric landing around 95 to 100%, gets it accepted.
- Distribution of Claude score on first attempt. The share of all submitted tasks that hit a 4 or 5 on the first attempt, our first-pass quality rate. It shows how consistently work lands right the first time instead of leaning on redo cycles. Goal: as high as possible.
- Redo rate. How many times a task gets sent back to redo_task before it moves through to audit. A task is redone when the review rating comes back at 3 or below, or the rubric eval flags items that may be wrong. Most tasks should settle within one iteration. Goal: as few as possible, ideally zero.
- Average handling time (AHT). The time between claiming a task from the pool and submitting it. Most tasks should take around 45 minutes, including any revisions, though this varies with complexity.
We aim for around 45 minutes per task on average, including any revisions and redos. Some tasks run longer and some run shorter depending on complexity. If a complex task needs closer to an hour to finish properly, that’s fine. Consistently running well over an hour is what we’ll follow up on.
New tasks drop in batches throughout the day, at no set time, so keep an eye on your Dashboard and claim them as they land. You can hold one task at a time. If you claim a task and don’t work on it within 24 hours, it detaches and returns to the pool for someone else, so only claim a task when you’re ready to start it.
To be confirmed.
Keep iterating first, since most tasks clear once you’ve addressed the real issues. But if the only flags left are hallucinations or oversensitive calls you disagree with, and the task is otherwise a solid 4 or 5, submit using the “Submit with Eval Override Request.” Add a few short bullets on why you think the flag is wrong.
Use this only when you truly believe a flag is incorrect, not as a shortcut past valid feedback.
Working location
Yes, you can work from any country.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
Get oriented, get connected, and start building chart tasks that a frontier model gets wrong.
Channels
New to the project? Start here.
Chartography announcements from the team.
Connect with other Chartography Contributors.
Discuss Chartography tasks and share tips.
Company-wide Workada announcements. Check daily before you task.
About the project
What Chartography is
Today’s frontier AI models read a simple bar chart without trouble. They fall apart on the charts professionals actually make decisions from: Kaplan-Meier survival curves, candlestick charts, contour maps, Sankey diagrams, Bode plots, phase diagrams, control charts, and weather soundings.
Chartography exists to find and document those failures. Each task pairs one complex chart with one hard question, one verified answer, and a clear step-by-step solution. Every task also has to clear the complexity bar, which means a frontier model has to get it wrong.
A brilliant question with a fuzzy answer isn’t usable. A rock-solid answer to a question the model already nails isn’t usable either. All five pieces together are what make a task deliverable.
What makes a strong chart task
- A complex visualization that rewards careful reading: multiple axes, overlapping series, dense or reused legend colors, a technical domain.
- A question with one definitive, checkable answer that takes several reasoning steps to reach.
- A clean step-by-step solution a reviewer could follow and independently verify.
- A model failure, documented, not asserted.
The bar: the model has to fail
Every task you create runs against a frontier model three independent times. A task is at an acceptable complexity level if two or more of those three attempts score as failures. If the model gets it right, the task isn’t hard enough yet, and that’s not a dead end: it’s your signal to tighten the question and try again.
Iterating on your own prompt until the model breaks is the core skill of this project, and it’s the part we most want your feedback on.
The domain field
Every task is tagged with a domain. Pick the closest fit; use Other if nothing matches.
- Finance and economics: Finance, Investing, Economics
- Science: General STEM, Chemistry, Biology, Physics, Geosciences
- Applied: Manufacturing, Supply Chain, Healthcare, Other
- Engineering: General Engineering, Mechanical, Electrical, Civil, Environmental
How to Build a Strong Task
A complex chart does a lot of the work for you, which makes this the place to spend time before you write anything. Whether you’re bringing your own chart or choosing one from the Chart Library, look for something busy and difficult to parse at first glance.
Strong charts usually have several of these features:
- Large numbers of overlapping data points
- Multiple axes with different units and scales
- Legends with many elements, especially ones that draw a distinction between similar colors or shapes
- Different types of data displayed together on the same plot
- Many clearly enumerated subdivisions
- Multiple curves that intersect at various points
- Colored areas that overlap in irregular ways
The list isn’t exhaustive, so if a chart makes you slow down and look twice, it’s probably a good candidate.
Understand the chart before you write. You either know what the chart shows out of the gate, or you spend time with the chart description and source URL to get up to speed. Either way, knowing what the data actually shows is what lets you write an informed, realistic prompt that stays unambiguous and points to a single defensible answer.
This guide covers what a strong prompt has to do, the seven guiding principles to check before you submit, and the question types that make the model reason instead of just count.
Read the Question Types & Prompt-Writing Guide before you write your next task, and keep it open while you work.
Prompt realism
The prompt should ask about the data underneath, not visual elements. Read the chart description in the claim sheet and open the source URL before you write. Then sense check: would a researcher ask this about this chart, and would the answer be a valuable insight in a professional setting?
Defensible, verifiable answers
If the scope allows several readings, it means several viable golden answers, and that’s a problem. Build the question on something that can’t be misinterpreted and layer guardrails into the prompt. Then sense check: would two reviewers following your solution land in the same place, and is anything about your golden answer debatable?
Every task you submit is reviewed against the quality rubric. Your reviewers work from that same document, so read it once before you start tasking and keep it open while you work.
If anything in this guide seems to disagree with the rubric, the rubric is the standard you’re graded on, so flag it and we’ll fix the guide.
- The chart made you slow down and look twice.
- One defensible answer, and you can say why.
- The answer comes from the image, not from general knowledge or a printed label.
- Scope names the region, the series, and the form of the answer.
- Your solution walks a reviewer to the answer, including the parts that are easy to miss.
- A researcher could plausibly ask your question.
Best ways to source your own chart for task acceptance
Chart choice shapes everything that comes after it. A chart that’s too simple limits how complex your prompt can get, no matter how much time you spend on it. This guide covers what to look for, where to look, and how to move from chart to prompt.
Seeded charts are available on your dashboard, and you can build strong prompts from those alone. Sourcing your own chart isn’t required.
That said, learning to source well gives you an option worth having. It lets you be proactive when tasking and bring in charts you already understand, in domains where you know the terminology and can spot the ambiguity faster.
The strongest charts share a few traits: multiple variables, overlapping series, or data points without clear labels, etc.. That ambiguity is what gives you room to build toward the more demanding question types instead of settling for a simple value extraction.
A chart that’s visually clean and easy to read at a glance usually works against you. There’s not enough there to construct a prompt that’s hard for the model to answer, and you’ll end up spending more time forcing complexity than the chart can support. If a chart takes you less than a minute to fully understand, it’s worth moving on.
Examples of charts:
Searching directly for charts, rather than starting from an article and hoping it contains one, tends to work best. You can keep a few charts in rotation at a time. If one isn’t yielding a strong task, you can switch to another instead of getting stuck reworking the task on a chart that AI is easily understanding.
Complexity only helps if the chart is actually legible. Before you settle on one, make sure the image is high enough resolution that axis labels, legends, and data points are clearly readable, not blurry, cropped, or compressed to the point of pixelation. Watch for watermarks or overlays that obscure part of the chart, and avoid screenshots where text has become distorted or cut off at the edges.
A chart that represents genuine data complexity is what you want. A chart that’s ambiguous because you can’t actually read it clearly is a different problem, and it will cause issues at review regardless of how strong the prompt is.
Before settling on a chart, read the article or the relevant section of the paper it comes from. This tells you whether there’s enough domain depth to support the question types that lean on interpretation or reasoning, and it surfaces the terminology you’ll want to use in your prompt. Charts backed by a source with real analytical depth consistently produce stronger prompts than charts pulled with no context.
Choose your question type after you’ve studied the chart, not before. Starting with “I want to write a calculation question” and then hunting for a chart that fits it usually leads to a forced prompt. Look at what the chart actually offers, and match it to one of the six question types:
- Value extraction: the chart lets you read a single value or range directly at a specified point, with little to no transformation needed.
- Comparison and ranking: the chart supports comparing multiple values, ordering them, or identifying an extremum, without needing a calculation to get there first.
- Calculation: the chart gives you clean numerical relationships to work with, so you can build a self-contained computation, a difference, ratio, growth rate, or similar.
- Enumeration: the chart has a clearly bounded set of items you can count or list against a stated condition, with unambiguous item boundaries. (Now Paused)
- Multi-step reasoning: the chart supports chaining two or more different kinds of operations together, where one step’s output feeds a different kind of step, like deriving a value and then comparing it.
- Conceptual interpretation: the chart requires understanding what it represents, not just reading it, whether that’s a qualitative judgment, identifying something from a signature, or locating a meaningful feature.
A chart with real complexity often supports more than one of these. Picking the type the chart is strongest for, rather than forcing a type it doesn’t support, is what separates a well-sourced chart from one that fights you.
- Does the chart have overlapping series, unlabeled points, or multiple variables?
- Would understanding it fully take you more than a minute?
- Does the source article or paper add real depth you can draw on?
- Have you matched the question type to what the chart actually supports, rather than the other way around?
If you can say yes to all four, you’re working from a chart that gives your prompt room to succeed.
Better ways to stump SOTA
Stumping the model is not about making the language of the prompt confusing or relying on clever wording. Rather, the difficulty should come from the reasoning required to solve the task. The strongest prompts contain a genuinely analytical step grounded in the chart, the domain, or both. That might involve interpreting an ambiguous pattern, connecting multiple variables, applying domain knowledge, performing a multi-step calculation, or deciding between competing conclusions using evidence from the chart.
Contributors with strong SOTA pass rates are generally not making their prompts harder to read. They are choosing charts and questions where the underlying reasoning is genuinely difficult, while keeping the prompt itself specific, detailed, and answerable.
Precision on the specific, not the general. The strongest prompts are often narrowly scoped. Instead of asking about the chart as a whole, direct the model to a specific part of the figure where careful reasoning is required. This could be a particular y-axis threshold, a defined region of the parameter space, a specific panel, or the behavior of a set of lines within a limited window.
Keep prompts short whenever possible, ideally one or two questions. Define the panel, time or value range, boundary rule, and whether the bounds are inclusive, then stop. Avoid adding extra layers that do not contribute to the core reasoning.
A strong model-stumping task usually starts with choosing the right chart. Focus on what the chart makes the viewer work to distinguish, not simply on the subject matter it covers.
Charts with overlapping series, unlabeled points, closely spaced values, multiple variables, crowded legends, or competing visual patterns tend to support harder questions than clean charts with a few clearly separated lines.
Before writing the prompt, spend a few minutes identifying where the chart itself requires careful interpretation. Look for regions where series converge or cross, values are difficult to distinguish, multiple conditions must be considered at once, or the same observation can be interpreted differently depending on context.
The goal is not to exploit poor readability, but to find a part of the figure where careful visual reasoning and domain understanding are genuinely required.
Beyond the general principle of finding ambiguity, these are seven recurring patterns worth checking for when picking a chart and a question.
Pattern 1: The sliver. The answer depends on something tiny: a thin line inside a bar, a near-zero bar next to tall ones, a 1% slice, a hairline ribbon, a tiny squeezed region. The model reads coarsely and misses anything a few pixels wide, even when a "none" or "zero" answer is correct. The model would rather find something than report nothing.
- Look for: a thin line in a bar, a near-zero bar, a 1% slice, a hairline ribbon, a tiny region
- Ask: count, which one, or is there any, so that missing the sliver changes the answer
Pattern 2: Shade discrimination. Two similar colours, and the answer depends on telling them apart: several steps of the same hue, adjacent similar hues, or a legend with more entries than you can hold in your head. The model lands one step off on a colour ramp and confidently matches the wrong legend entry. The colour is the category here, not decoration.
- Look for: 6 or more legend colours, a colour ramp with 4 or more steps, similar neighbouring hues
- Ask: which category or which band, named exactly as the key names it
Pattern 3: The decoy. The right answer isn't where your eye goes first. The model goes for whatever is biggest, steepest, or most obvious, so if the correct answer sits in a squashed or off-to-the-side part of the chart while something large and obvious is wrong, it takes the bait. Reading conventions make great decoys: box top vs. whisker end, local vs. global peak, a log or flipped axis.
- Look for: a squashed region near an axis, local vs. global maximum, box top vs. whisker, a log or inverted axis
- Ask: largest, first, or highest, defined precisely, where the obvious answer is wrong
Pattern 4: Two-key filter. Each condition is easy alone. Both at once is not. Colour and shape, series and panel, a threshold on x and on y: holding two visual conditions across many candidates makes the model drop items that qualify and admit ones that don't. Per-panel legends are especially effective, since the model reads the wrong panel's key.
- Look for: colour and shape legends, per-panel legends, thresholds on both axes
- Ask: which items meet both conditions, with a small answer set
Pattern 5: Near-ties and hairline crossings. The model rounds small differences away: two bars almost the same height, a curve poking a hair above another, two close crossings. This has become the most popular pattern, which is a problem: on multi-line charts it can turn into a tolerance argument rather than a real stump. Use it only where the gap is genuinely visible.
- Look for: almost-equal bars, hairline crossings, two crossings in a narrow window
- Ask: strictly higher or lower within a stated window, or how many crossings between a and b
Pattern 6: Boundary gap. Ask for the width of a band, not the position of its edge: the thickness of a stacked slice or ring, not a drawn boundary. The chart never prints that number, so the model reads the edge, or the running total, instead of the individual gap. The bands can be perfectly ordinary in size. This isn't about something small. It's about measuring the wrong thing.
- Look for: stacked areas or bars, ribbon or band charts, radial stacks, anything drawn as a thickness
- Ask: how thick is this band at x, where the outer edge peaks somewhere else entirely
Pattern 7: Occlusion recovery. The data is there, just hidden behind something else: a marker or line segment covered by another mark. Distinct from Pattern 1 (small but visible) and Pattern 6 (genuinely absent), this is plotted but hidden, and the answer has to be reconstructed from what's on either side. Never tell the model the value is hidden. That collapses it into an ordinary read.
- Look for: crowded time series with overlapping lines, stacked scatter markers, later series drawn over earlier ones
- Ask: what is X's value at this point, where X is buried under other series there
For all seven: the pattern tells you which chart to pick and where the hard part is. It doesn't change how you word the question. That still has to be the ordinary thing a professional in the field would ask, in the chart's own vocabulary. A question that only makes sense as a way of catching the model out will be sent back as unrealistic.
Before working on the prompt, take time to review the source URL from which the chart was derived. This will give you the broader context needed to understand what the chart is measuring, what each variable represents, and what relationships or trends the figure is intended to highlight.
Reviewing the source also helps you use the correct domain-specific terminology and frame the task in a way that reflects how a professional in that field would reasonably interpret the figure. That context should inform the prompt, Golden Answer, and Step-by-Step solution so that all three are conceptually accurate, technically precise, and grounded in the source material rather than relying only on surface-level visual reading.
Ask the question a professional would ask. Focus on three things: what the underlying data represents, how the figure contributes to the broader analysis or argument, and what someone working in that field would realistically want to learn from it.
Use that context to frame the prompt around a meaningful analytical objective rather than a surface-level visual task. The question should reflect how a domain professional would interpret, compare, or use the data, while the Golden Answer and Step-by-Step solution should apply the same terminology and reasoning consistently.
Of the six question types outlined in the Chartography Question Types and Writing Guide, prompts that combine conceptual interpretation with multi-step reasoning are generally the most effective at challenging the model. By contrast, simple value extraction is usually the easiest question type for the model to answer correctly.
If your prompts are passing too easily, check whether they are asking the model only to read values from the chart rather than reason with them. Stronger prompts should require the model to connect multiple observations, perform calculations or comparisons, and interpret what those results mean in the context of the chart.
Reaching that level of depth often requires going beyond the figure itself. Read the source article or paper, understand what the variables represent, and spend a few minutes looking up unfamiliar terminology. That context makes it much easier to create a question that reflects genuine domain reasoning rather than surface-level chart reading.
Difficulty should come from the reasoning required, not from complicated wording. Ask a question that a professional in the field could reasonably ask, and make sure two careful reviewers looking at the same evidence would arrive at the same answer.
Sometimes a prompt fails to challenge the model because the chart itself does not contain enough complexity to support a difficult, meaningful question. Precise wording cannot fully compensate for a figure with limited data, relationships, or analytical depth.
If you have already tried clearly defining the scope, using appropriate domain-specific terminology, and combining conceptual interpretation with multi-step reasoning, but the task still passes too easily, the chart itself may be the limiting factor.
At that point, avoid forcing complexity through convoluted wording, arbitrary conditions, or unnecessary question chaining. It is usually more effective to select a different chart that naturally supports deeper comparison, calculation, and interpretation than to continue reworking a fundamentally simple figure.
- Have I reviewed the source URL and understood what the chart is actually measuring, why it matters, and how someone in the domain would use it?
- Would a professional in this field reasonably ask this question, or does it feel like a visual scavenger hunt created only to stump the model?
- Am I using the chart's domain-specific terminology rather than generic language?
- Does the task require conceptual interpretation, calculation, comparison, or multi-step reasoning rather than simply reading a value, colour, label, or count?
- Does the chart itself contain enough complexity to support the question, such as overlapping series, unlabeled observations, multiple variables, crossings, thresholds, or competing patterns?
- Am I using the chart's underlying data as part of the reasoning rather than asking only about visual characteristics?
- Is the prompt short and focused, ideally one or two connected questions, rather than a chain of unrelated tasks?
- Does each question build naturally on the previous one, rather than question hopping between unrelated parts of the figure?
Examples of Stellar Work
Browse examples by domain: click a tab below to switch.
1. Food, alcohol, & tobacco
2. Non-energy industrial goods
3. Services
4. All-items
5. Energy
From July 2016 to July 2026, Non-energy industrial goods remained relatively closest to 0% inflation, indicating the most stable inflation rate around the zero-inflation benchmark among the components listed.
January 1st 2012: second index is TLT, and third index is SPY.
There is no information to answer the question.
Next, we need to identify the horizontal axis on the bottom part of the chart, which is labeled “Time”, and that has markers in the following points in time: from left to right, January 1st 2010, January 1st 2011, January 1st 2012 and January 1st 2013.
Having identified the different points in time, we now focus on the two dates that are specified in the prompt: January 1st 2011 and January 1st 2012. Starting with the first date, we need to check the relative vertical position of each index price line at that point over the horizontal axis. Maintaining our position over that point in time, we check the position of each index price line from top to bottom. The indexes appear as follows: RWR, EEM, SPY, GSP, TLT.
Moving on to the second date, we repeat the process, checking the relative vertical position of each index price line at that point over the horizontal axis. The indexes appear as follows: RWR, TLT, SPY, GSP, EEM.
Having checked the relative position of all the index price lines at both points in time, we can now answer the question: for January 1st 2011, the second index is EEM, and third index is SPY; for January 1st 2012, second index is TLT, and third index is SPY.
To answer the third question we need to locate the COY price index line at both points in time. However, we can see that the chart does not show any data points for COY index. Therefore, we can conclude that there is no information to answer the question.
Going back to the outermost end of the spiral, we can notice that the thin stripes making up the outermost strata of the spiral represent the presence of different organisms, each one labeled accordingly with a different color. Multicellular life is the 5th band from the outside going in, and it is colored in a forest green shade. If we trace this line back to the inside of the spiral, we can notice that it overlaps with two different eons, the phanerozoic and the proterozoic.
Notice that glaciations are represented as blue segments within the innermost tier of the spiral. Since the spiral represents a linear timeline, the arc length of each segment is directly proportional to the length of the event it represents. Therefore, the longest glaciation is represented by the longest blue glaciation segment, which is clearly the pongola glaciation.
Determine the site type; an LR-Market is a triangle. Look for dark purple triangles. There is exactly one purple triangle on the plot on the left. There are 8 field sites represented by dark purple dots. There is only 1 LR market, so we have to find what color the LR market is attributed to in panel B. It should be green with the caption "unknown".
Spectral Width = 0nm ; (0,60)
Spectral Width = 0.1nm ; (20,60)
Spectral Width = 0.3nm ; (30,60)
Spectral Width = 0.5nm ; (40,50)
Question 2: For each spectral width chart, find the smallest window of Y-values that will include all the graphed data. You are only allowed to choose values of Y that are labeled on the Major gridlines however (multiples of 10, from 0 to 100), so for example if the lowest data point is 15%, the lower bound of the data will have to be 10%. Similarly, if the highest value is 55%, the top bound will have to be 60%. The fact that the bounds are inclusive means that they include their own values in the reported range, e.g., if the curve's maximum is touching but not crossing 50%, then 50% would be an appropriate upper bound.
Chartography FAQ
Training and Calibration
Training and calibration have a combined paid time cap of 1.5 hours for Chartography. This limit has been set to cover everything you need to complete both and we're confident it provides enough time for most contributors. Once you reach the cap, you're welcome to continue with your timer turned off, but any time beyond that won't be compensated.
Your Pod Lead reviews your training and you are notified once they've gone over it. Once you pass, you move on to calibration.
Calibration is a task you complete once you've passed training. You get two attempts to pass it. If your first attempt doesn't pass, it comes back to you to revise and try again. Your Pod Lead reviews your calibration task, so if you've been waiting a while, check in with them for the current timeline in your pod.
Writing strong tasks
It means a frontier model can already answer your question reliably, so the task isn't usable as is. The bar in Chartography is that the model has to fail. If a task comes back "Within SOTA capability," iterate to make it harder: more multi-step reasoning, more disambiguation, fewer direct reads.
This one trips people up because it sounds negative, but it's good news. It means the model failed to answer your question, which is the goal. If a task instead comes back as "Within SOTA capability" (you may also see "previous attempt" or "draft"), that means your question wasn't strong enough yet.
Skip a chart when you can't make it both single-answer/verifiable and hard enough, even after trying the usual difficulty levers: multi-hop reasoning, miscounting traps, overlaps, legend or color traps, scale changes, absence answers, off-chart label reading. Mark a chart "not suitable for task creation" when the chart itself is the problem: too trivial, answer labels already printed on it, illegible or cropped, or it requires outside data to solve.
Plan for roughly an hour on average, including revisions and resubmission. Some tasks take longer if you need extra iterations to stump the model.
Yes, as long as the prompt still has exactly one defensible answer (for example, "it never exceeds that value"). Two things to watch: don't word it so it sounds like you're assuming the event happens, and make the "none/never" outcome an explicitly acceptable answer.
Yes, as long as there's one defensible answer that's answerable from the chart alone, and the count isn't trivial or already labeled. Strong counting tasks usually involve real visual reasoning, like overlapping marks, ambiguous grouping, or legend/color categories, rather than just reading a printed number.
Reviewers grade against the Chartography Quality Rubric, across these dimensions:
- Chart Quality
- Prompt Realism
- Reasoning Requirement
- Self-Containment
- Defensible and Verifiable Answer
- Answer Correctness
Reviews and feedback
Bring it to your Pod Lead with the task link and the specific rubric dimension you think was mis-scored.
There's no standard appeal form. Take the same route as above: bring it to your Pod Lead with the task link and the dimension in question.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
Everything for the Video Caption project: your Slack channels, what the project is about, and how to complete each section of a captioning task.
Channels
New to the project? Start here for setup, guidelines, and getting started.
Project announcements from the Video Caption team.
Connect with other Video Caption Contributors.
Discuss Video Caption tasks and share tips.
Your pod-specific channel. Your Pod Lead will add you when you join the project.
Company-wide Workada announcements. Check daily before you task.
Project FAQ
Training and calibration together are capped at 5 hours.
Average handle time is 1.5 to 2 hours, depending on the complexity of the task.
There are no location restrictions. You can work from anywhere.
We're always looking to improve this guide. If you'd like to see something added, let your Pod Lead know.
Everything for the Remedy project: your Slack channels, what the project is about, and how to complete a task.
Channels
New to the project? Start here for setup, guidelines, and getting started.
Project announcements from the Remedy team.
Connect with other Remedy Contributors.
Discuss Remedy tasks and get help from the team, we're online most of the day.
Company-wide Workada announcements. Check daily before you task.
About the project
Project overview coming soon
This section will cover what the Remedy project is, what a task looks like, and what good work looks like. Send the project overview and training document to your Workada contact and it will be added here.
Project FAQ
Can't find what you need? Post in #remedy-discussions, we're online most of the day.
Getting started
You'll need a login on workada.com (use the email tied to your interview) and access to GPT 5.6 Sol on High Reasoning. If you don't already have a GPT Plus subscription, ask your project lead. Workada covers a month of it and arranges it directly with you.
From your dashboard, open Tasks in the left sidebar, then claim "Create Remedy Task" under "Available to Claim." That opens the intake form where you'll submit your scrubbed input files, the source paper link, and your prompt.
The Project Remedy Instructions doc covers the workflow end to end. Read it in full before your onboarding session, and revisit it alongside the Prompt Writing and Golden Response guidelines as you go.
Writing prompts
The best prompts integrate multiple pieces of experimental evidence the way a domain expert actually reasons through a paper, not a single lookup. Two patterns reviewers consistently point to:
- Ask the model to combine results across several figures, assays, or data files to reach one conclusion, rather than reading a single value off one source.
- Give the model a real decision point, one where a plausible but wrong analytical route exists, and specify a fallback like NOT_IDENTIFIABLE for when the data doesn't support a clear answer.
Once a task has a well-formed prompt, the rest of the task tends to follow. Full guidance and examples live in the Prompt Writing Guidelines.
Specific enough that there's exactly one defensible way to do the analysis. Reviewers will flag a prompt if there are multiple reasonable interpretations. For example, several valid ways to define a fold-change calculation, which time points to require, or whether to rank by FDR or raw p-value all count as ambiguity. If more than one defensible answer exists, tighten the prompt until only one does.
That's still useful data. Log it in the Prompts With No Model Failure sheet:
- Add a new row with the prompt, the model's response, and a shareable chat link.
- You can leave the notes column blank for now.
- Copy the task ID from the three-dot menu on the task and paste it into column A.
If the model seems close to failing, keep the same task open and try a revised prompt. If it's not close, it's usually faster to start a new task.
Scrubbing & data prep
Models are surprisingly resourceful about extracting an answer from metadata rather than the analysis itself. One task looked like a clean model failure until the reviewer noticed the file name itself contained the answer. Once it was renamed to something opaque, the model actually started failing as intended. Treat file names, sample labels, and folder structure as part of the prompt.
Yes. The task form has fields for both, plus a short text field explaining what you scrubbed. Reviewers use the original-to-scrubbed comparison to spot lingering answer-revealing details and to build better guidance on what to scrub next time.
Renaming files and sample identifiers has resolved this for other contributors, and it's often enough on its own to stop the model from recognizing the paper. If it persists, share the files you're using with your project lead in #remedy-discussions and they'll help identify what's leaking.
Working with the model
Yes, you can do this in the same conversation you'll submit. Just make sure the prompt and response you paste into the task are the ones from the first turn, not the debugging exchange that follows.
No. With memory off, a new task starts clean even if the earlier task hasn't been deleted. This is the recommended way to work through multiple prompts drawn from the same publication without cross-contamination.
Paste the visible response and include the shareable chat link. Hidden portions, like reasoning traces or "thought for X seconds" content, don't need to be pasted in for now; the chat link covers that.
Reviews & feedback
The [Remedy] Attempted Tasks and Reviews Tracking sheet is the source of truth for every task's status and score. It shows whether a task is good to go, needs changes, or needs a new sub-decision, and reviewer notes on what to fix live in column J.
- Read the feedback in column J of the tracking sheet.
- Questions about it go in #remedy-discussions, where each task gets its own thread.
- Make the update in your Workada Dashboard, and confirm the model still fails after the change.
- Once resubmitted, mark column M, "Updated By Attempter."
Reviewers do another pass from there, and this can take more than one round for a task to land.
The Golden Response Guidelines are a companion doc to the Prompt Writing Guidelines, clarifying exactly what belongs in the golden response field, added after reviewers noticed contributors interpreting that field differently. Worth reading in full before you submit your next task.
It means there's more than one defensible way to reach an answer, so the task doesn't have a single objectively correct result. The fix is usually to pin down the exact analytical decisions in the prompt itself: which comparisons to make, how to define the metric, which statistical threshold to use, and so on. See "Writing prompts" above.
Pay & referrals
Payouts run weekly. Your GPT Plus reimbursement is added to the payout following your subscription, not the same one, so expect it to show up a few days after you first mention the expense to your project lead.
Yes, referrals are welcome. You'll receive $150 per referral for anyone who completes an accepted task. Reach out to your project lead directly to make the introduction.
Support
Log it in the [Remedy Sites] Being Blocked by Workada Timer sheet and message your project lead directly. They'll get the timer updated as soon as possible.
Yes. Drop-in Zoom office hours run periodically, no need to stay for the whole session. Bring your task and your questions, get aligned with a reviewer, and head out once you're unblocked. Announcements for these go out in #remedy-announcements.
We're always looking to improve this guide. If you'd like to see something added, let your Pod Lead know.
Coming soon
Resources for the Multimodal project are on the way. Your Pod Lead will let you know as soon as this page is ready.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
Coming soon
Resources for the Lotus project are on the way. Your Pod Lead will let you know as soon as this page is ready.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
Coming soon
Resources for the Multi-hop Reasoning project are on the way. Your Pod Lead will let you know as soon as this page is ready.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
Coming soon
Resources for the Finance Acquisition project are on the way. Your Pod Lead will let you know as soon as this page is ready.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
Coming soon
Resources for the Sheets Acquisition project are on the way. Your Pod Lead will let you know as soon as this page is ready.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
Coming soon
Resources for the Sheets Artifact Collection project are on the way. Your Pod Lead will let you know as soon as this page is ready.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.
These are the standards for using Workada’s Dashboard and Slack. In short: use the platform only for legitimate task work, keep one secure account, submit accurate work, keep everything you access confidential, and treat everyone with respect. Read the full Code of Conduct →
On the Dashboard
Account integrity
- Keep only one account. Don’t create another after a suspension or termination without written permission.
- Keep your login credentials secure, and never share, sell, or transfer access to your account.
- Make sure your personal and payment information is accurate, current, and in your own name.
- Complete any identity-verification request from Workada right away.
Performing tasks
- Take on only tasks you can complete accurately and on time.
- Follow the scope, specifications, and instructions provided with each task.
- Don’t submit work that’s fabricated, copied from unauthorized sources, or done by someone else in your place.
- Don’t try to game quality-control or verification. Repeated inaccurate work can lead to removal from projects or account deactivation.
Confidentiality and data
- Treat all Workada materials, customer data, and task content as confidential.
- Don’t download, copy, or store confidential information beyond what a task requires.
- Don’t share task details or customer data with third parties, on social media, or anywhere else.
- When a task ends or your account closes, delete all confidential information from your devices and cloud storage.
Intellectual property
- Everything you create through the Dashboard belongs to Workada. Don’t keep copies or try to use, sell, or license it on your own.
- Don’t submit content that infringes anyone else’s intellectual property.
- Don’t use Workada’s materials for anything outside your tasks.
Software and systems
- Use only approved browsers and apps. Don’t reverse-engineer, decompile, or extract source code from any Workada software or extension.
- Don’t probe, scan, or test the security of the Dashboard or any connected system.
- Don’t introduce malware, scripts, or bots that manipulate the platform or task results.
- Report any security issue or suspected breach to support@workada.co right away.
Compliance
- Use the Dashboard only for lawful purposes, following all applicable laws.
- Don’t work from, or on behalf of, any country or party under applicable sanctions or export controls.
- Don’t upload or submit content that’s illegal, discriminatory, harassing, or otherwise harmful.
Workada can investigate suspected breaches, hold payment during an investigation, remove you from projects, or suspend or deactivate your account, at its reasonable discretion under the Terms of Use.
On Slack
Professional standards
- Communicate respectfully at all times. Personal attacks, insults, and hostile behavior aren’t allowed.
- No harassment, discrimination, or bullying on the basis of race, gender, age, religion, disability, national origin, sexual orientation, or any other protected characteristic.
- Handle disagreements constructively, and escalate anything unresolved to support@workada.co rather than letting it become a public argument.
- Keep your language inclusive and appropriate for a professional, international audience.
Permitted use
- Use Slack to discuss Workada tasks, workflows, platform updates, and professional development.
- Post in the right channel, and keep messages concise. Off-topic or repetitive posts may be removed.
- Don’t use Slack to solicit business, advertise third-party services, or recruit Contributors away from Workada.
Confidentiality and content
- Don’t share confidential information, task details, customer data, materials, or work product, in any channel, DM, or thread.
- Don’t post screenshots, documents, or files from the Dashboard unless Workada authorizes it.
- Don’t post anything illegal, defamatory, obscene, threatening, or harmful, and don’t post spam or phishing links.
- Don’t share or endorse misinformation about Workada, its customers, or other Contributors.
Privacy and moderation
- Don’t share copyrighted material or third-party personal data without authorization.
- Don’t record, screenshot, or distribute private conversations without everyone’s consent.
- Moderators may monitor channels and remove content that breaks these standards. Report violations to support@workada.co.
- Repeated or serious violations can lead to removal from the workspace, suspension, or account deactivation.
Running into a snag mid-task is frustrating: we get it, and we’re always working to make Workada better. Here’s how our support process works, plus a few tips that help us take care of you faster.
How to submit a support ticket
- You run into a problem while tasking, or you have a pay or timer issue.
- Open the Support module in the Workada Dashboard and email support@op.workada.co, or email that address directly using the email linked to your Workada dashboard.
- In the message, include your name, your email address, the Task ID (if needed), and a brief description of your issue.
- You’ll get an automatic reply confirming that a member of the support team will get back to you within the next 2 business days. Want to add more detail on something you’ve already submitted? Reply directly on the original ticket, don’t open a new one.
- Workada Support follows up on the same email thread to resolve the issue: either guiding you through a fix or resolving it directly on Workada’s end.
- No response after 5 days? Reach out to your Pod Lead and give them the ticket number, they can escalate your unresolved ticket.
Do’s and Don’ts at a Glance
From your application to your first task: here’s the full path, stage by stage, so you always know where you are and what comes next.
Team & Community
Your Slack home base. Join these channels, introduce yourself, and check in regularly.
Sign Up
- Apply. Submit your application to Workada.
- Resume screen. Our team reviews your application and resume.
- Interview invitation. If you pass the screen, you’ll receive an email inviting you to an interview.
- Interview. You’ll meet with a member of our team over Zoom.
Onboarding
- Get access to Slack.
- Get your Pod assignment (more on this below).
- Complete Persona verification.
- Set up your bank account: allow extra time for this step; it’s the one that most often takes a few tries.
- Set up your timer.
Pod Assignment
Every Contributor is assigned to a Pod Lead: your go-to person for questions, guidance, and escalations.
- Slack channel. You’re added to your Pod Lead’s Slack channel.
- Welcome email. Your Pod Lead sends you a welcome email to get you started.
Training
- Finish training.
If you passYou get access to calibration tasks.If you don’t pass yetReach out to your Pod Lead: it depends on the situation, and in some cases additional resources and training are available to help you get there.
- Calibration tasks. Complete your calibration tasks to fine-tune your ratings before real work begins.
If you passYou’ll be able to continue with your project-specific tasks.If you don’t passReach out to your Pod Lead: it depends on the project, in some projects additional resources and training are available to help you get there.
Not passing training and calibration may lead to offboarding.
Project Assignment
- Start tasking. You’re fully set up: browse available tasks and jump in whenever you’re ready. If you need help at any point, reach out to your Pod Lead.
We’re always looking to improve this guide. If you’d like to see something added, let your Pod Lead know.