CHEON:AI

Intelligence with a human temperature

CHEON:AI builds search foundation models — rerankers, embeddings and Text2Cypher — for enterprise search and RAG. Open weights on Hugging Face, API access for production.

Most search stops at matching words. Ours keeps going — retrieving by meaning, reranking by relevance, and asking the knowledge graph for exact answers.

01 / Rerank

Reranker

The precision stage of retrieval. A reranker reads the query together with the candidates your search returned and reorders them, so the passages that actually answer the question come first — an upgrade for the search you already run, with no re-indexing.

Released cheon-reranker-0.6b-v1

02 / Retrieve

Embeddings

The first stage of search. An embedding model turns documents and queries into vectors, so nearest-neighbour search finds what a query means — not only the words it happens to use.

In development

03 / Query

Text2Cypher

For questions that live in relationships. Text2Cypher writes the Cypher query a graph database runs, so answers come back as exact records rather than passages that merely look similar.

In development

Latest release

cheon-reranker-0.6b-v1

A reranker for legal and administrative document search, built on Qwen3-0.6B.

Task
Text ranking
Parameters
609M
Base model
Qwen/Qwen3-0.6B
Languages
Korean
License
cc-by-nc-4.0
Max sequence length
12,288 tokens
Max candidates per context
50
Updated
Sep 29, 2026
import torch
from transformers import AutoModel, AutoTokenizer

repo = "cheonai/cheon-reranker-0.6b-v1"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, torch_dtype=torch.float32)
model.eval()

# Query + candidates go into one sequence. Track each candidate's token span
# as you build it instead of hardcoding positions. Candidates don't need to
# share a language - the tokenizer and span logic don't care.
query = "query: 이 사건의 처리 기한은 언제까지인가요?\n"
candidates = [
    "candidate (en): The appeal must be filed within 30 days of the decision.",
    "candidate (ko): 이의신청은 처분을 안 날부터 30일 이내에 제기해야 합니다.",
    "candidate (zh): 上诉必须在裁定后30天内提出。",
    "candidate (ja): 不服申立ては、決定を知った日から30日以内に行う必要があります。",
]

ids = tok(query, add_special_tokens=False)["input_ids"]
doc_spans = []
for text in candidates:
    piece = tok(text, add_special_tokens=False)["input_ids"]
    start = len(ids)
    ids += piece
    doc_spans.append((start, len(ids)))

input_ids = torch.tensor([ids])
attention_mask = torch.ones_like(input_ids)
corpus_ids = [0] * len(candidates)  # optional per-candidate corpus/collection tag

with torch.no_grad():
    scores = model(
        input_ids=input_ids,
        attention_mask=attention_mask,
        doc_spans=doc_spans,
        corpus_ids=corpus_ids,
    )
print(scores)  # higher = more relevant

Access

Two ways to use the models.

Open weights

Weights, code and model cards are on Hugging Face, released under each model's license. Load them with transformers and run them on your own hardware.

Browse on Hugging Face

API and commercial use

Production traffic and commercial licensing are arranged with us directly. Tell us about your corpus, languages and query volume.

Request API access

Our name, our approach

체온 (che-on) —
body temperature.

Che-on is the warmth of a living body — 36.5 degrees. It is our standard for machine intelligence: always on, never cold.

The colon in our name is the pause between a question and its answer. That pause is where we work.

The CHEON:AI wordmark — a bold CHEON, a quieter AI, and a copper colon marking the pause between question and answer.

For your product

Put our models to work

Request API access

From the lab

Notes on search and RAG

Read the blog

FAQ

Questions, answered.

What is a search foundation model?
A model built for the retrieval side of search and RAG: finding candidates, putting the best of them first, and querying structured knowledge. These models decide what an LLM gets to read before it answers.
Which model should I start with?
If you already run search, start with a reranker. It reorders the results your current system returns, so it needs no new index and shows its effect on the queries you already have.
How are the models licensed?
Each model carries its own license, stated on its model card: cheon-reranker-0.6b-v1 (cc-by-nc-4.0). If a license does not cover your use, ask us about commercial licensing.
Can I run the models on my own hardware?
Yes. The weights are published on Hugging Face and load with transformers; each model card has the exact code.
How do I get API access?
Send a request through the contact form with your use case, the data you search over and your expected query volume.

Start a conversation

Let's raise the temperature
of your search.

Whether you are building retrieval from scratch or tuning a RAG system that already runs, we would like to hear about it.

Get in touch →