Your advertising privacy

With your permission, we use the Meta Pixel to measure our Facebook and Instagram ads and improve who sees them. Meta receives page visits, completed enquiry events, your IP address and browser information, and may set advertising cookies. Your form answers are not sent. How Meta uses data

We remember this choice for 180 days. Meta's first-party advertising cookies can last up to 90 days. Rejecting stops future events and removes those cookies here; it does not erase data already received by Meta. Change your choice at any time in Privacy settings. Proptonomy privacy policy

All news
Product

Proptonomy runs on the best AI models.

Inside the Proptonomy benchmark: 23 AI models tested on 67 real scenarios from nine operators, so the model behind your guests is chosen by evidence. GPT-6 Luna wins; the Chinese models trail.

By Mathias Haugsbø, Co-founder and CTO

Scatter chart of 23 AI models on the Proptonomy benchmark: judge accuracy against cost per task on a logarithmic scale. GPT-6 Luna sits top left at 64% for a quarter of a cent; GPT-6 Astra and the Claude models sit top right at up to 84% for 50 cents or more; the Chinese models sit in the middle and bottom.

Proptonomy runs day-to-day operations for short-term rental managers: it answers guests, coordinates cleaners, chases handymen and keeps the admin team and owners informed. All of these operations are processed by large language models. Which model we put behind the wheel is therefore the single biggest quality decision we make, and we must base it on real benchmarking.

So we built our own benchmark. This post is what it is, what it found on 1 October 2026, and why we keep running it.

Not customer-service chit-chat

Public benchmarks test maths, code and trivia. Very few test whether a model can run a rental operation. So we built a benchmark from 67 real situations, frozen from production at the exact moment our agent had to decide what to do, across operators on three continents. Names and contact details are anonymized, everything else is as it happened.

Most of the scenarios come from internal staff operations, not from simple guest questions. A few of them, in short-hand:

  • Check when the guest will check out, adjust the cleaning schedule, and inform the cleaner if anything changed.
  • A partner confirms hot water is restored. Close the right task and tell the guest, without disturbing the admin team.
  • An admin asks which active listings are off the new 16:00 check-in and 10:00 check-out standard, across the whole portfolio.
  • Staff suggest moving a same-day guest to a sibling unit that is already booked. The right answer is to say no, and why.

We collected scenarios across maintenance, internal operations, bookings, check-in and check-out. They arrive in English, German, Spanish, Korean, Chinese and Norwegian, and many of them are graded as difficult. We oversample the hard ones on purpose: an easy benchmark tells you nothing.

How a model is scored

Each model gets exactly the tools our Proptonomy agent has. It can look up bookings, tasks, properties and the knowledge base, and every lookup returns the real data we recorded at that moment. If the model "sends a message" or "creates a task", that happens in a sandbox copy. No real guest, cleaner or owner is ever contacted.

A scenario only counts as passed when all three checks say yes:

  • Did the right thing. The system ends in the right state: the right task created, the right person messaged, nothing forbidden done.
  • Told the right person. Whoever needed the key information actually received it.
  • A judge agrees. A second AI reads everything the model did and confirms it is correct, grounded in the facts, with no made-up claims.

The score is the share of scenarios passed. This is the same method as the public τ²-bench, so the numbers are comparable in kind to what the model vendors publish, except that the tasks are ours. Hard scenarios are run three times, because a model that gets it right once in three is not one you can put in front of a guest.

The results, 1 October 2026

We have now run 23 models through the benchmark. The chart above puts all of them on one picture: accuracy against cost per task. The matched run we did this week, with every model graded on all 67 scenarios by the same judge, came out like this:

  • GPT-6.1 Sol: 79% accuracy, 4.6 cents per task, 27 seconds per task.
  • Claude Sonnet 5.5: 76% accuracy, 9.2 cents per task, 9 seconds per task.
  • GPT-6 Luna: 64% accuracy, a quarter of a cent per task, 10 seconds per task.

And the overall winner is GPT-6 Luna. Not because it is the most accurate, it is not. GPT-6 Astra scores highest of everything we have tested, at around 84%, and Sol and Sonnet beat Luna by a clear margin. But Luna costs 18 times less than Sol and 37 times less than Sonnet per task, at the same speed, and it sits on the cost frontier with Sol. Per dollar spent it returns more than sixteen times the score of its nearest rival. For an operation that makes thousands of these decisions a day, that is the model that runs most of Proptonomy.

The second finding is the one in the headline. The Chinese models are falling behind. DeepSeek v4.1 Flash is the best of them at around 68%, which is respectable, but it takes 113 seconds per task, and while it costs less on paper, in reality it costs more than ten times what Luna does. Below it the field thins out fast: Qwen 3.8 Flash and GLM 5.3 Flash in the mid-fifties, MiMo v2.6 Pro and MiniMax M3 around the halfway mark, and Qwen 3.7 Flash, hy3 and MiMo v2.5 between 13% and 30%. On a task where the wrong answer means a guest locked out or a cleaner sent to the wrong apartment, that is not a gap we can route around with prompting.

A caveat we hold ourselves to: these are curated hard decision points, not a random sample of traffic, so the percentages are not a production success rate. They are a ranking under identical conditions, and the ranking is what we act on.

Why we keep running it

New models ship every few weeks. Each one arrives with a vendor chart and a claim. Within a day of a release we can put it through the same scenarios and know, in our own terms, whether it would run a portfolio better than what we have. Nothing changes in production on the strength of a benchmark alone, but nothing changes without one either.

That matters because of who is on the other end. Operators run their properties on Proptonomy in Norway, the United Kingdom, Ireland, Switzerland, Italy, South Africa, New Zealand, Mexico and the United States. Our own sister company Heimby manages around 240 properties across seven Norwegian cities. In New Zealand the agent coordinates close to eighty cleans a day for one of Australasia's largest operators, from Wellington to Christchurch and Tekapo. In London, Zurich, Davos and Dublin it talks to guests and contractors in their own language. The scenarios come from that footprint, which is why a model that aces an English guest FAQ and fumbles a Korean pre-booking inquiry or a Spanish message from on-site staff does not win here.

We will publish the next round when the next generation of models lands. If you run short-term rentals and want to see what the current winner does with your operation, send us a message.

Give your operations team the night off.

Tell us where to reply. We come back with the same thirty-day figures for your portfolio, and a thirty-minute walk through the desk if you want one.

Pick a time that suits you. We only show times we are free for. Sending without one is fine too.

Checking the calendar…

Built by the team behind Heimby, after three years running one of Norway's largest short-term rental operations.

No commitment