prashanth-ds-ml Claude Sonnet 5 commited on
Commit
e3dc7eb
·
1 Parent(s): ba625ed

Add HF Spaces Docker deployment: transliteration endpoint, static UI serving

Browse files

server/transliterate.py wraps IndicXlit (ai4bharat-transliteration),
imported defensively in app.py since fairseq only installs cleanly on
Linux - local Windows dev keeps working without it, the Space container
has it. New GET / serves the UI directly instead of requiring a separate
static server. ui/rachana-desk.html's API_BASE now self-detects file://
vs served-origin, and the Roman-script notice gets a real Transliterate
action instead of a static message. Dockerfile/requirements-deploy.txt/
README.md frontmatter configure the Docker SDK Space.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

.dockerignore ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ .git/
2
+ .gitattributes
3
+ .gitignore
4
+ .env
5
+ __pycache__/
6
+ *.pyc
7
+ .venv/
8
+ .venv-unsloth/
9
+ wandb/
10
+ unsloth_compiled_cache/
11
+ data/
12
+ models/*
13
+ !models/exp3_gemma3_4b_public_v0_1/
14
+ models/exp3_gemma3_4b_public_v0_1/checkpoint-*/
15
+ .obsidian/
16
+ .agents/
17
+ .claude/
18
+ reports/
19
+ scripts/
20
+ IndicXlit_python_inference.ipynb
21
+ AUDIT_REPORT.md
Dockerfile ADDED
@@ -0,0 +1,45 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Rachana News - HF Spaces deployment image.
2
+ # Python 3.10 is required, not just conventional: fairseq's config
3
+ # dataclasses crash on Python 3.11+'s stricter mutable-default rule
4
+ # (see errors.md 2026-07-21/22).
5
+ FROM python:3.10-slim
6
+
7
+ WORKDIR /app
8
+
9
+ # build-essential: fairseq compiles a small C extension (fairseq.libbleu)
10
+ # at install time - this is exactly what native Windows lacked (no MSVC),
11
+ # and what WSL2/Linux ships by default via gcc/g++.
12
+ RUN apt-get update && apt-get install -y --no-install-recommends \
13
+ build-essential git \
14
+ && rm -rf /var/lib/apt/lists/*
15
+
16
+ COPY requirements-deploy.txt .
17
+
18
+ # Standard PyPI torch build - the special cu128/Blackwell pin this project
19
+ # needs for local dev (RTX 5050) does not apply to Spaces GPU hardware
20
+ # (T4/A10G/etc, all standard architectures with kernels in the default
21
+ # distribution already).
22
+ RUN pip install --no-cache-dir torch && \
23
+ pip install --no-cache-dir -r requirements-deploy.txt
24
+
25
+ # fairseq/ai4bharat-transliteration: pip<24.1 works around a real
26
+ # omegaconf metadata-rejection bug (a malformed PyYAML version specifier
27
+ # in omegaconf's own wheel, rejected by pip>=24.1's stricter parsing).
28
+ # Cloning fairseq manually (rather than letting pip clone it internally)
29
+ # avoids an intermittent git smart-protocol error seen during this
30
+ # project's WSL2 validation. Both fixes are logged in errors.md
31
+ # 2026-07-21/22, where this exact sequence was proven working.
32
+ RUN pip install --no-cache-dir "pip<24.1" && \
33
+ git clone --depth 1 https://github.com/facebookresearch/fairseq.git /tmp/fairseq_src && \
34
+ pip install --no-cache-dir -e /tmp/fairseq_src && \
35
+ pip install --no-cache-dir ai4bharat-transliteration
36
+
37
+ COPY server/ server/
38
+ COPY ui/ ui/
39
+ COPY models/exp3_gemma3_4b_public_v0_1/ models/exp3_gemma3_4b_public_v0_1/
40
+
41
+ ENV PYTHONUNBUFFERED=1
42
+ # HF Spaces' Docker SDK expects the app to listen on 7860 by default.
43
+ EXPOSE 7860
44
+
45
+ CMD ["uvicorn", "server.app:app", "--host", "0.0.0.0", "--port", "7860"]
README.md CHANGED
@@ -1,63 +1,44 @@
 
 
 
 
 
 
 
 
 
 
1
  # Rachana News
2
 
3
- A Telugu-language AI assistant for newsroom editors and reporters, built by fine-tuning Gemma 3 (4B) into a domain specialist for editorial tasks: drafting, paraphrasing, editing, translation, and transliteration.
4
 
5
  **Full product spec:** See `Rachana_News_PRD.md`
 
6
 
7
- ## Repo Structure
8
 
9
- - `scripts/` — Python pipeline for evaluation
10
- - `prepare_eval_data.py` — samples ~10 prompts per capability from public datasets
11
- - `run_eval.py` — runs prompts through local Gemma 3 4B via Ollama
12
- - `score_eval.py` — scores capabilities with automated metrics or outputs manual-rating CSVs
13
- - `data/processed/` — evaluation data
14
- - `eval_prompts.jsonl` — input prompts across 8 capabilities
15
- - `eval_results.jsonl` — model outputs
16
- - `manual_ratings_*.csv` — three CSVs awaiting human ratings:
17
- - `manual_ratings_paraphrase.csv`
18
- - `manual_ratings_fluency.csv`
19
- - `manual_ratings_notes_to_article.csv`
20
- - `reports/` — findings
21
- - `baseline_v0.md` — baseline eval of unmodified Gemma 3 4B-IT; answers PRD §11 open question #1
22
- - `.obsidian/` — Obsidian vault config (project is indexed as a vault)
23
 
24
- ## Setup
25
 
26
- Requires Python 3.10, `pip`, and Ollama (installed separately).
27
 
28
  ```bash
29
  python -m venv .venv
30
  source .venv/bin/activate # or .venv\Scripts\activate on Windows
31
- pip install -r requirements.txt
32
- ollama pull gemma3:4b
 
33
  ```
34
 
35
- ## Running the Pipeline
36
-
37
- ```bash
38
- # Step 1: Prepare evaluation data (samples from public datasets)
39
- python scripts/prepare_eval_data.py
40
-
41
- # Step 2: Run Gemma 3 4B on each prompt via local Ollama
42
- python scripts/run_eval.py
43
-
44
- # Step 3: Score outputs with automated metrics, emit manual-rating CSVs for tasks without ground truth
45
- python scripts/score_eval.py
46
- ```
47
-
48
- **Awaiting human review:** The three `manual_ratings_*.csv` files need 1–5 ratings (fluency/register/fact-preservation/task-success columns) filled in by hand; no automated scoring exists for paraphrase, fluency, or notes→article capability.
49
-
50
- ## Key Findings
51
-
52
- **Baseline evaluation (Gemma 3 4B-IT, unmodified)** is complete — see `reports/baseline_v0.md` for full results.
53
 
54
- **Open question #1 (§11 of PRD):** Is continued pretraining needed?
55
- **Answer:** Tentatively no. Base Telugu fluency and grounded generation are solid already; targeted fine-tuning on weak capabilities (headline generation coherence, summarization register, transliteration task data) is more cost-effective than a pretraining stage.
56
 
57
- Real weak points identified:
58
- - **Headline generation**: incoherent word insertions, hallucination at word level
59
- - **Summarization**: accurate but wrong register (too formal, not terse)
60
- - **Transliteration**: exact-match score understates real quality due to romanization-convention variance; needs dedicated task data but metric caveat is important
61
- - **Translation (En→Te)**: solid starting point, lower sample size (n=5), worth rechecking
62
 
63
- **Next steps:** See PRD §11 items 2–5 (human-written seed set size, external teacher model decision, Phase 3 scoping, Rachana Data Studio extension for preference collection).
 
1
+ ---
2
+ title: Rachana News
3
+ emoji: 📰
4
+ colorFrom: blue
5
+ colorTo: yellow
6
+ sdk: docker
7
+ app_port: 7860
8
+ pinned: false
9
+ ---
10
+
11
  # Rachana News
12
 
13
+ A Telugu-language AI writing assistant for newsroom editors and reporters, built by fine-tuning Gemma 3 (4B) into a domain specialist for editorial tasks: headline generation, paraphrasing, grammar correction, translation, transliteration, and research.
14
 
15
  **Full product spec:** See `Rachana_News_PRD.md`
16
+ **Current project state:** See `progress.md` (rewritten each session — the authoritative status snapshot)
17
 
18
+ ## What's here
19
 
20
+ - `server/` — FastAPI backend serving the exp3 QLoRA adapter (headline/grammar/paraphrase/translate/summarize/transliterate) plus a controlled-RAG Research Assistant over curated Telugu news feeds
21
+ - `ui/rachana-desk.html` — the writer workspace UI, served by the backend at `/`
22
+ - `models/exp3_gemma3_4b_public_v0_1/` — the trained LoRA adapter (base model `google/gemma-3-4b-it` downloads at container startup, gated - requires `HF_TOKEN`)
23
+ - `scripts/` — the training/eval pipeline (dataset prep, QLoRA training, adapter evaluation) - not part of the deployed image, local-dev only
24
+ - `Rachana_News_PRD.md`, `progress.md`, `decisions.md`, `errors.md`, `data-sources.md` — project memory; read `progress.md` first for current status
 
 
 
 
 
 
 
 
 
25
 
26
+ ## Running locally
27
 
28
+ Requires Python 3.10.
29
 
30
  ```bash
31
  python -m venv .venv
32
  source .venv/bin/activate # or .venv\Scripts\activate on Windows
33
+ pip install -r requirements-deploy.txt
34
+ pip install torch --index-url https://download.pytorch.org/whl/cu128 # or the CUDA build matching your GPU
35
+ uvicorn server.app:app --reload --port 8000
36
  ```
37
 
38
+ Set `HF_TOKEN` (accepted `google/gemma-3-4b-it` license required) in a local `.env` file — never commit it.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
39
 
40
+ Transliteration (`/api/transliterate`) needs `fairseq`/`ai4bharat-transliteration`, which only install cleanly on Linux (see `errors.md` 2026-07-21/22) - the endpoint degrades gracefully (`available: false`) when those packages aren't present, so local Windows dev works without them.
 
41
 
42
+ ## Deploying
 
 
 
 
43
 
44
+ See `Dockerfile` — builds a self-contained image (server + UI + transliteration) for HF Spaces' Docker SDK. Requires a GPU-tier Space (bitsandbytes 4-bit quantization needs CUDA) and `HF_TOKEN` set as a Space repository secret.
requirements-deploy.txt ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Serving-only dependencies for the deployed Space - trimmed from the
2
+ # training-oriented requirements.txt. Excludes datasets/sacrebleu/
3
+ # rouge-score/jiwer/wandb/trl/pandas, which server/app.py, server/
4
+ # guardrails.py, and server/research.py never import.
5
+ #
6
+ # torch and fairseq/ai4bharat-transliteration are installed separately in
7
+ # the Dockerfile (see decisions.md 2026-07-22 for why: torch needs a CUDA
8
+ # index URL pip can't express in a plain requirements file, and fairseq
9
+ # needs the pip<24.1 + build-tooling sequence validated in WSL2 this
10
+ # session).
11
+ fastapi==0.139.2
12
+ uvicorn==0.51.0
13
+ feedparser==6.0.12
14
+ python-dotenv==1.2.2
15
+ requests==2.34.2
16
+ transformers==5.14.1
17
+ accelerate==1.14.0
18
+ peft==0.19.1
19
+ bitsandbytes==0.49.2
server/app.py CHANGED
@@ -23,6 +23,7 @@ import torch
23
  from dotenv import load_dotenv
24
  from fastapi import FastAPI
25
  from fastapi.middleware.cors import CORSMiddleware
 
26
  from peft import PeftModel
27
  from pydantic import BaseModel
28
  from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
@@ -30,6 +31,18 @@ from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
30
  from server.guardrails import check_fact_preservation, check_headline_needs_review, detect_script
31
  from server.research import fetch_all_entries, search_entries, summarize_article
32
 
 
 
 
 
 
 
 
 
 
 
 
 
33
  ROOT = Path(__file__).resolve().parent.parent
34
  load_dotenv(ROOT / ".env")
35
 
@@ -92,6 +105,8 @@ app.add_middleware(
92
  @app.on_event("startup")
93
  def _startup():
94
  load_model()
 
 
95
 
96
 
97
  class TextRequest(BaseModel):
@@ -193,6 +208,18 @@ def api_research(req: ResearchRequest):
193
  return {"results": results}
194
 
195
 
 
 
 
 
 
 
 
 
 
 
 
 
196
  @app.post("/api/feedback")
197
  def api_feedback(req: FeedbackRequest):
198
  FEEDBACK_LOG.parent.mkdir(parents=True, exist_ok=True)
 
23
  from dotenv import load_dotenv
24
  from fastapi import FastAPI
25
  from fastapi.middleware.cors import CORSMiddleware
26
+ from fastapi.responses import FileResponse
27
  from peft import PeftModel
28
  from pydantic import BaseModel
29
  from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
 
31
  from server.guardrails import check_fact_preservation, check_headline_needs_review, detect_script
32
  from server.research import fetch_all_entries, search_entries, summarize_article
33
 
34
+ try:
35
+ # fairseq/ai4bharat-transliteration only install cleanly in a Linux
36
+ # environment (see errors.md 2026-07-21/22 - native Windows needs both
37
+ # Developer Mode and an MSVC compiler, neither pursued for local dev).
38
+ # Import defensively so the rest of the server still runs locally on
39
+ # Windows without this dependency; the deploy container has it.
40
+ from server.transliterate import transliterate_text
41
+ from server.transliterate import load_engine as load_transliteration_engine
42
+ TRANSLITERATION_AVAILABLE = True
43
+ except ImportError:
44
+ TRANSLITERATION_AVAILABLE = False
45
+
46
  ROOT = Path(__file__).resolve().parent.parent
47
  load_dotenv(ROOT / ".env")
48
 
 
105
  @app.on_event("startup")
106
  def _startup():
107
  load_model()
108
+ if TRANSLITERATION_AVAILABLE:
109
+ load_transliteration_engine()
110
 
111
 
112
  class TextRequest(BaseModel):
 
208
  return {"results": results}
209
 
210
 
211
+ @app.post("/api/transliterate")
212
+ def api_transliterate(req: TextRequest):
213
+ if not TRANSLITERATION_AVAILABLE:
214
+ return {"available": False, "words": []}
215
+ return {"available": True, "words": transliterate_text(req.text)}
216
+
217
+
218
+ @app.get("/")
219
+ def ui_index():
220
+ return FileResponse(ROOT / "ui" / "rachana-desk.html")
221
+
222
+
223
  @app.post("/api/feedback")
224
  def api_feedback(req: FeedbackRequest):
225
  FEEDBACK_LOG.parent.mkdir(parents=True, exist_ok=True)
server/transliterate.py ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ Roman-to-Telugu transliteration via IndicXlit (ai4bharat-transliteration).
3
+
4
+ Not usable on native Windows (fairseq needs a symlink at build time that
5
+ requires Developer Mode, plus an MSVC compiler for its C extensions) - see
6
+ errors.md 2026-07-21/22. Works cleanly in a Linux container, where this
7
+ module is actually loaded, matching the WSL2 validation done this session.
8
+ """
9
+ from ai4bharat.transliteration import XlitEngine
10
+
11
+ _engine = None
12
+
13
+
14
+ def load_engine():
15
+ global _engine
16
+ if _engine is None:
17
+ print("[transliterate] loading XlitEngine('te') ...")
18
+ _engine = XlitEngine("te", beam_width=10, rescore=False)
19
+ print("[transliterate] engine ready")
20
+ return _engine
21
+
22
+
23
+ def transliterate_text(text: str, topk: int = 3) -> list[dict]:
24
+ """Transliterate each whitespace-delimited word; punctuation passes through unchanged."""
25
+ engine = load_engine()
26
+ results = []
27
+ for word in text.split():
28
+ core = word.strip(".,!?;:\"'()[]")
29
+ if not core:
30
+ results.append({"input": word, "candidates": [word]})
31
+ continue
32
+ candidates = engine.translit_word(core, topk=topk)["te"]
33
+ results.append({"input": word, "candidates": candidates})
34
+ return results
ui/rachana-desk.html CHANGED
@@ -789,7 +789,11 @@
789
  (function () {
790
  "use strict";
791
 
792
- var API_BASE = "http://127.0.0.1:8000";
 
 
 
 
793
 
794
  // ============ Task definitions (Quick Tools) ============
795
  // Live-wired day-1 tasks. notes->article and transliteration are
@@ -1053,7 +1057,9 @@
1053
  html += '<textarea class="input-area" id="input-area">' + escapeHtml(ts_.input) + '</textarea>';
1054
  if (SCRIPT_SENSITIVE_TASKS[task.id]) {
1055
  html += '<div class="review-banner" id="script-notice" style="display:none">' +
1056
- '⚠ ఈ ఇన్‌పుట్ రోమన్/మిక్స్డ్ లిపిలో ఉంది. ఉత్తమ ఫలితాల కోసం తెలుగు లిపిలో టైప్ చేయండి లేదా అతికించండి.</div>';
 
 
1057
  }
1058
  var genLabel = task.render === "research" ? "శోధించండి · Search" : "రూపొందించండి · Generate";
1059
  var genLoadingLabel = task.render === "research" ? "శోధిస్తోంది…" : "రూపొందిస్తోంది…";
@@ -1156,6 +1162,31 @@
1156
  updateScriptNotice(inputArea.value);
1157
  });
1158
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1159
  Array.prototype.forEach.call(document.querySelectorAll("[data-dir]"), function (btn) {
1160
  btn.addEventListener("click", function () {
1161
  ts_.direction = btn.getAttribute("data-dir");
 
789
  (function () {
790
  "use strict";
791
 
792
+ // Local double-click-the-file workflow needs an absolute origin (no
793
+ // origin to resolve a relative fetch against under file://); served from
794
+ // the FastAPI app itself (local dev via uvicorn, or the deployed
795
+ // container) same-origin relative paths just work.
796
+ var API_BASE = (location.protocol === "file:") ? "http://127.0.0.1:8000" : "";
797
 
798
  // ============ Task definitions (Quick Tools) ============
799
  // Live-wired day-1 tasks. notes->article and transliteration are
 
1057
  html += '<textarea class="input-area" id="input-area">' + escapeHtml(ts_.input) + '</textarea>';
1058
  if (SCRIPT_SENSITIVE_TASKS[task.id]) {
1059
  html += '<div class="review-banner" id="script-notice" style="display:none">' +
1060
+ '⚠ ఈ ఇన్‌పుట్ రోమన్/మిక్స్డ్ లిపిలో ఉంది. ఉత్తమ ఫలితాల కోసం తెలుగు లిపిలో టైప్ చేయండి లేదా అతికించండి. ' +
1061
+ '<button class="action-btn" id="transliterate-btn" type="button">తెలుగులోకి మార్చండి · Transliterate</button>' +
1062
+ '</div>';
1063
  }
1064
  var genLabel = task.render === "research" ? "శోధించండి · Search" : "రూపొందించండి · Generate";
1065
  var genLoadingLabel = task.render === "research" ? "శోధిస్తోంది…" : "రూపొందిస్తోంది…";
 
1162
  updateScriptNotice(inputArea.value);
1163
  });
1164
  }
1165
+ var translitBtn = document.getElementById("transliterate-btn");
1166
+ if (translitBtn) {
1167
+ translitBtn.addEventListener("click", function () {
1168
+ translitBtn.disabled = true;
1169
+ translitBtn.textContent = "మారుస్తోంది…";
1170
+ fetch(API_BASE + "/api/transliterate", {
1171
+ method: "POST", headers: { "Content-Type": "application/json" },
1172
+ body: JSON.stringify({ text: inputArea.value })
1173
+ }).then(function (r) { return r.json(); }).then(function (data) {
1174
+ if (!data.available) {
1175
+ flash("Transliteration is not available in this environment.");
1176
+ return;
1177
+ }
1178
+ var converted = data.words.map(function (w) { return w.candidates[0]; }).join(" ");
1179
+ inputArea.value = converted;
1180
+ ts_.input = converted;
1181
+ updateScriptNotice(converted);
1182
+ }).catch(function (e) {
1183
+ flash("Transliteration failed: " + e.message);
1184
+ }).finally(function () {
1185
+ translitBtn.disabled = false;
1186
+ translitBtn.textContent = "తెలుగులోకి మార్చండి · Transliterate";
1187
+ });
1188
+ });
1189
+ }
1190
  Array.prototype.forEach.call(document.querySelectorAll("[data-dir]"), function (btn) {
1191
  btn.addEventListener("click", function () {
1192
  ts_.direction = btn.getAttribute("data-dir");