Ferhany

Evaluating language models in the real world.

Ferhany is an independent research lab in Istanbul. We build benchmarks, agentic evaluations and training environments for the everyday tasks people bring to language models, in many languages.

Our position

Most benchmarks measure how well a model writes code. Most people use models for something else.

They ask, compare, buy, plan and play, often in a language other than English. We study models in that wider world, score them on outcomes rather than impressions, and turn each evaluation into an environment that can train the next model.

Research

Four lines of work, one standard: the result must be something we can check.

  • Language games

    Taboo, Codenames, word puzzles and riddles, built natively in each language. They test communication, reasoning and culture, and every game ends in an outcome code can verify.

    First release
  • Real-world agentic tasks

    Multi-step tasks from everyday life, such as shopping, travel and customer service, with simulated users and real tool constraints. Scored on whether the task was done.

    In development
  • Humans and LLM judges

    When does an LLM judge agree with people? We compare native-speaker ratings with model judgments across languages and tasks, and measure where judges drift.

    In development
  • Training environments

    Each evaluation, packaged as an environment for reinforcement learning with verifiable rewards, ready for post-training.

    Next

An example

What a single test looks like.

Describe çay without saying:

demlikbardakRizekahveşeker

Turkish adds suffixes, so “demliğe” counts as “demlik”. The check has to understand morphology, not just match strings.

Checked automatically

Forbidden words
none, in any form
Guesser model
names the word
Turns used
counted

Contact

[email protected]

For model builders who need evaluation beyond code, and researchers working on judges, agents or training environments.