How the check works
When you press the button, your description goes to Jev, a decision model built for software rather than chat. It does not write a reply. It answers a fixed list of questions about your task, and for each question it gives a probability for every possible answer:
- How much an agent could do on its own: none of it, a small part, most of it with your approval, or all of it end to end.
- How often it happens and how long one round takes by hand.
- What kind of work it is, from invoice chasing and reporting to support, research and data entry.
- Which apps it touches, such as email, spreadsheets, a CRM or a help desk.
- What could go wrong: whether a mistake would be harmless or serious, whether it needs someone on site, and whether it handles other people's personal or financial data.
The verdict at the top comes from the first question. The bars under it show the full spread, so you can see when the model is torn between two answers instead of trusting a single label.
How the cost estimate is worked out
The calculator starts from the model's guess at how often the task happens and how long it takes, and you can change both. Then:
- Hours a month = runs a month × minutes each time ÷ 60. Every working day counts as 21 runs, every week as 4.33.
- Cost by hand = hours a month × your hourly cost.
- Hours an agent could take over = hours a month × the share an agent can do, less a tenth of that for checking its work.
The share comes from the model's probabilities: none counts as 0%, a small part as 25%, most of it as 75% and all of it as 100%, weighted by how likely each answer is. It is an estimate for deciding where to start, not a quote.
Which tasks AI agents handle well in 2026
The tasks that score highest share four traits. They start on a schedule or after a clear event, like a new email, order or form. They live in software you can connect. The steps are the same each time, even when the content changes. And a mistake is easy to spot and fix. Weekly reports, invoice reminders, inbox sorting, lead follow-up and meeting notes fit that shape, which is why they come up again and again in the ROI case studies.
Tasks score low when they need hands and eyes in the physical world, when every case is different, or when one wrong move costs real money or trust. Those can still get help, for example an agent that prepares a draft or gathers the facts, but a person should stay in charge.
What the checker cannot see
It only knows what you write. It cannot tell whether your apps allow outside access, whether your data is tidy, or whether your team will trust the result. Treat a high score as a reason to try, and a low score as a prompt to describe the task in more detail or split it into smaller steps. Splitting often helps: "send the weekly report" may score better on its own than "run the weekly sales meeting".
Questions
How does the checker decide whether AI can do a task?
Your description goes to Jev, a decision model from TypeSafe, with a fixed set of questions: how much of the task an agent could do on its own, how often it happens, how long it takes, what kind of work it is, which apps it touches, and whether it needs a person on site or handles sensitive data. Jev returns a probability for every possible answer, and the page shows those probabilities instead of a single guess.
Is the result accurate?
It is a fast first opinion, not an audit. The model only knows what you wrote, so a one-line description gets a rougher answer than a description that names the apps, the steps and how often the task happens. Each answer shows how sure the model is, and you can correct the frequency and time before trusting the cost figures.
How is the monthly cost of my time worked out?
Runs a month times minutes per run gives hours a month. Hours times your hourly cost gives what the task costs by hand. The hours an agent could take over are the hours a month times the share of the task the model thinks an agent can do, less about a tenth of that for checking its work.
What hourly cost should I use?
Use what an hour of the person doing the task costs the business: salary plus overheads divided by working hours, or your own billable rate if you run the business. The calculator starts at $30 an hour so you can see a result straight away.
Do you store what I type?
The description is sent to the model to answer the questions and is not saved by Gravity, unless you press Share it or ask us to set it up as an agent; a task shared that way stays private and is never published. We count how many times the tool runs each day, nothing more. Avoid pasting names, account numbers or other private details; the checker does not need them.
What can I do with a task that scores high?
Tasks that start on a schedule or after a clear event, live in software, and have a low cost of mistakes are the easiest to hand to an AI agent. The guides linked under each result show how that kind of task is usually automated, and Gravity runs that kind of recurring task as an agent: it is in private alpha and you can apply on the alpha page.