Evaluating language models in the real world.
Ferhany is an independent research lab in Istanbul. We build benchmarks, agentic evaluations and training environments for the everyday tasks people bring to language models, in many languages.
Our position
Most benchmarks measure how well a model writes code. Most people use models for something else.
They ask, compare, buy, plan and play, often in a language other than English. We study models in that wider world, score them on outcomes rather than impressions, and turn each evaluation into an environment that can train the next model.
Research
Four lines of work, one standard: the result must be something we can check.
-
Language games
Taboo, Codenames, word puzzles and riddles, built natively in each language. They test communication, reasoning and culture, and every game ends in an outcome code can verify.
First release -
Real-world agentic tasks
Multi-step tasks from everyday life, such as shopping, travel and customer service, with simulated users and real tool constraints. Scored on whether the task was done.
In development -
Humans and LLM judges
When does an LLM judge agree with people? We compare native-speaker ratings with model judgments across languages and tasks, and measure where judges drift.
In development -
Training environments
Each evaluation, packaged as an environment for reinforcement learning with verifiable rewards, ready for post-training.
Next
An example
What a single test looks like.
Describe çay without saying:
demlikbardakRizekahveşeker
Turkish adds suffixes, so “demliğe” counts as “demlik”. The check has to understand morphology, not just match strings.
Checked automatically
- Forbidden words
- none, in any form
- Guesser model
- names the word
- Turns used
- counted
Two replies to the same Turkish customer complaint. Which one would the customer prefer? Native-speaker raters and LLM judges answer the same question, blind to each other and to which model wrote each reply.
Measured
- Agreement
- judge vs. people
- Position bias
- order swapped
- Length bias
- controlled
- By language
- and by task
Contact
For model builders who need evaluation beyond code, and researchers working on judges, agents or training environments.