Open weights
Weights, code and model cards are on Hugging Face, released under each model's license. Load them with transformers and run them on your own hardware.
Browse on Hugging FaceCHEON:AI builds search foundation models — rerankers, embeddings and Text2Cypher — for enterprise search and RAG. Open weights on Hugging Face, API access for production.
Most search stops at matching words. Ours keeps going — retrieving by meaning, reranking by relevance, and asking the knowledge graph for exact answers.
The retrieval stack
Dense vectors find candidates by meaning, not only by matching words.
In development
02 · RerankEvery candidate is read against the query, and the strongest evidence moves to the top.
Released cheon-reranker-0.6b-v1
03 · QueryQuestions about entities and relationships become Cypher, answered by the graph itself.
In development
The model writes from the best evidence it was given, and can cite it.
The precision stage of retrieval. A reranker reads the query together with the candidates your search returned and reorders them, so the passages that actually answer the question come first — an upgrade for the search you already run, with no re-indexing.
Released cheon-reranker-0.6b-v1
The first stage of search. An embedding model turns documents and queries into vectors, so nearest-neighbour search finds what a query means — not only the words it happens to use.
In development
For questions that live in relationships. Text2Cypher writes the Cypher query a graph database runs, so answers come back as exact records rather than passages that merely look similar.
In development
Latest release
A reranker for legal and administrative document search, built on Qwen3-0.6B.
import torch
from transformers import AutoModel, AutoTokenizer
repo = "cheonai/cheon-reranker-0.6b-v1"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, torch_dtype=torch.float32)
model.eval()
# Query + candidates go into one sequence. Track each candidate's token span
# as you build it instead of hardcoding positions. Candidates don't need to
# share a language - the tokenizer and span logic don't care.
query = "query: 이 사건의 처리 기한은 언제까지인가요?\n"
candidates = [
"candidate (en): The appeal must be filed within 30 days of the decision.",
"candidate (ko): 이의신청은 처분을 안 날부터 30일 이내에 제기해야 합니다.",
"candidate (zh): 上诉必须在裁定后30天内提出。",
"candidate (ja): 不服申立ては、決定を知った日から30日以内に行う必要があります。",
]
ids = tok(query, add_special_tokens=False)["input_ids"]
doc_spans = []
for text in candidates:
piece = tok(text, add_special_tokens=False)["input_ids"]
start = len(ids)
ids += piece
doc_spans.append((start, len(ids)))
input_ids = torch.tensor([ids])
attention_mask = torch.ones_like(input_ids)
corpus_ids = [0] * len(candidates) # optional per-candidate corpus/collection tag
with torch.no_grad():
scores = model(
input_ids=input_ids,
attention_mask=attention_mask,
doc_spans=doc_spans,
corpus_ids=corpus_ids,
)
print(scores) # higher = more relevant
Access
Weights, code and model cards are on Hugging Face, released under each model's license. Load them with transformers and run them on your own hardware.
Browse on Hugging FaceProduction traffic and commercial licensing are arranged with us directly. Tell us about your corpus, languages and query volume.
Request API accessOur name, our approach
Che-on is the warmth of a living body — 36.5 degrees. It is our standard for machine intelligence: always on, never cold.
The colon in our name is the pause between a question and its answer. That pause is where we work.
FAQ
Start a conversation
Whether you are building retrieval from scratch or tuning a RAG system that already runs, we would like to hear about it.
Get in touch →