HunyuanOCR on GX10: a licence that excludes the EU
HunyuanOCR is a small (1 billion parameters) model from Tencent that reads documents and returns where on the page each line is — coordinates. It is technically interesting. But its licence says it does not apply in the European Union — which includes Bulgaria. This lesson shows how to read a licence before installing anything, and what to use instead.
01What you'll learn
- Why a licence is read before the installation, not after it.
- What exactly the HunyuanOCR licence says about the territory and where in the text you find it.
- What HunyuanOCR is: size, tasks, and what "coordinates" (spotting) means.
- Which other restrictions the licence has and why they matter to a company.
- What to use instead in the EU — OCR with coordinates, under a licence without a territorial exclusion.
02Before you start
- Nothing to install. The lesson is for reading and one optional terminal check — a machine of the NVIDIA GB10 class (for example ASUS Ascent GX10 or DGX Spark) is not needed.
- Know where the model would run and who would use it: the country of the machine and of the users. The Territory in the licence is measured by where it is used.
- If it is for a client — know which side is the client and which is yours: the licence also governs the use of the model's outputs.
03Steps
-
The licence first
On the model's page on Hugging Face the "License" field shows
tencent-hunyuan-community— not Apache-2.0 and not MIT, but Tencent's own licence. The text itself is in theLICENSEfile. You can also view it from the terminal:bash · on the machinecurl -sL https://huggingface.co/tencent/HunyuanOCR/raw/main/LICENSE | head -n 3 curl -sL https://huggingface.co/tencent/HunyuanOCR/raw/main/LICENSE | grep -n -i "territory"The first command shows the title and the sentence about the EU. The second shows every line with the word "territory" — so you find the territory clauses without reading 20 pages at once.
-
What the licence says about the territory
Five places in the text say the same thing:
Where What it says (meaning) The first line The agreement does not apply in the European Union, the United Kingdom and South Korea and is limited to the "Territory". Definition 1(l) "Territory" = the whole world, excluding the EU, the United Kingdom and South Korea. Section 2 The right to use, copy, modify and distribute is granted for the Territory only. Section 5(c) Do not use, copy, modify, distribute or display the materials or their outputs outside the Territory; such use is "unlicensed and unauthorized". Acceptable Use Policy, item 1 Do not use the model or its derivatives outside the Territory. Two things that are often missed: (1) the restriction applies to the output too (the recognised text, the JSON, the Markdown) — not only to the weights themselves; (2) "I only downloaded the weights from Hugging Face" changes nothing — the terms are about use, not about where it was downloaded.
✅RuleIf any part of the use — the machine, the users or the client — is in the EU, skip the model and pick an alternative. Do not look for a "workaround". -
The other restrictions in the licence
Even outside the EU the licence is not "free". In short, what else it says:
- Do not improve another AI model with the model or its outputs (section 5b) — for example, do not train your own model on text recognised with it.
- Acceptable Use Policy: a ban on military use, on high-stakes automated decisions (health, employment, credit, law enforcement and others) and on publishing machine-generated content without a clear note that it is machine-generated.
- If you offer a service built on it: include a copy of the licence and a "Notice" file, state clearly who the real provider is and that Tencent is not affiliated with the service (sections 3d and 3e).
- Over 100 million monthly active users — a separate licence from Tencent is needed (section 4).
- Law and courts: the laws and courts of Hong Kong (section 9).
This is our reading of the published text, not legal advice. For real company use — a lawyer.
-
What the model is
According to the model page (version 1.5, announced on 07.07.2026) HunyuanOCR is a "lightweight, end-to-end OCR-specialized vision-language model": you give it an image and an instruction, you get text back. One model does four things — document parsing, text spotting with coordinates, information extraction and translation of text in images.
What According to the documentation Size about 1 billion parameters (1.12 billion per Hugging Face), BF16 format Versions 1.5 sits at the root of the repository; 1.0 is archived in the v1.0/folderNew in 1.5 faster decoding (DFlash), running through llama.cpp on a laptop or an ordinary graphics card, images up to 4K, context up to 128K Languages For 1.0 the documentation claims "over 100 languages". Bulgarian is not named in the texts we checked — ⚠️ not confirmed. Translation of text in images is described for 14 languages (German, Spanish, Turkish, Italian, Russian, French, Portuguese, Arabic, Thai, Vietnamese, Indonesian, Malay, Japanese, Korean) — Bulgarian is not among them Running vLLM (OpenAI-compatible server), "transformers" (5.13.0 or newer), llama.cpp (GGUF) Hardware For 1.0: Linux, an NVIDIA GPU, ~20 GB of GPU memory with vLLM. For 1.5: CUDA 13 What "coordinates" (spotting) are: for each line of text the model returns not only the letters but also where in the picture the line is — a box. This is useful when you want to find a field in a form or place something at an exact spot in a PDF. The official client has separate task types for this (
spotting_jsonandspotting_hunyuan) and for document parsing (doc_parse), tables, formulas and charts.ℹ️The vendor's numbersTencent publishes comparison tables in which the model leads — on its own in-house text-spotting test and on OmniDocBench. These are the vendor's numbers; we have not reproduced them. OCR benchmarks are crowded with entries; a trial on your own pages decides. -
Does it work on GB10 — honestly: we do not know
Since 24.07.2026 both the Hugging Face page and the GitHub README describe one shared environment for running it: CUDA 13, Python 3.12, vLLM 0.25.1 or newer and flash-attn 2.8.3 built by hand. The first 1.5 release used three separate environments that cannot be mixed; they remain in the documentation as "lighter recipes" for machines without CUDA 13 — for example vLLM 0.18.1 with CUDA 12. None of the texts says anything about Arm64 (the GB10 processor is not x86). Have we run it — no. And after the previous step the question is moot for the EU anyway.
-
What to use instead in the EU
If you need line coordinates, there are tools with a lesson here. We checked the licences of PaddleOCR (Apache-2.0) and dots.mocr in their own lessons; for the others, check the licence the same way as in step 1. The full comparison is in lesson 04-216.
Tool What for Lesson PaddleOCR 3.x 🔒 local Apache-2.0 licence. The documented result fields include dt_polys(the line boxes). Bulgarian withlang="bg". PP-StructureV3 returns Markdown and tables04-215 Tesseract 🔒 local Classic OCR on the CPU; Bulgarian with bul; word positions throughimage_to_dataor TSV/hOCR output04-213 EasyOCR 🔒 local Returns a box, the text and a confidence score for each line; Bulgarian with bg04-214 dots.mocr 🔒 local A whole-page model: box, category and text for each element (a block, not a line). MIT plus a supplementary agreement with no EU exclusion, but with bans on sensitive personal data; Bulgarian is not named 04-211 Before you take any of them into a product — repeat step 1 for it: open the licence and check the terms. That habit is the point of this lesson.
-
Security — the minimum
- Tencent's online demo is a Tencent service — documents you upload there leave your machine. Do not use it for confidential documents.
- Do not put other people's real documents into tests and lessons; use your own or invented ones.
- Recognised amounts, dates and names are a draft — before they enter accounting, a contract or a database, a person reviews them.
04Check
- Find and read the first line of the licence and the "Territory" clause.
- You can say in one sentence why HunyuanOCR is not used in Bulgaria.
- You know the restriction applies to the outputs, not only to the weights.
- You have an alternative with coordinates chosen and know which lesson it is in.
- For every new model you will open the licence first.
Quiz
1. What is the "Territory" according to the HunyuanOCR licence?
2. A Bulgarian company wants to put HunyuanOCR into a service for clients in Varna. What follows from the licence?
3. What can you use in the EU if you need line coordinates?
4. Which other restriction is in the HunyuanOCR licence?
05What's next
06Sources
- HunyuanOCR on Hugging Face 🌐 global — the model page (1.5), licence field, size, ways to run it.
- Tencent Hunyuan Community License Agreement — the first line, definition 1(l), sections 2, 3, 4, 5, 9 and the Acceptable Use Policy.
- HunyuanOCR on GitHub — README of 1.5 (news of 07.07.2026, environments) · README of 1.0 (languages, spotting prompt, requirements).
- HunyuanOCR-1.5 (arXiv 2607.04884) · 1.0 technical report (arXiv 2511.19575).