Current release: v0.5.0
Contact: sjran@cnu.edu.cn
Paper: arXiv:2608.29612
Repository: github.com/ranshiju/Ran-ASKS
Project updates: CHANGELOG.md
General-reader introduction (Chinese): Read on GitHub | Download or print the PDF | Version scope
The Markdown edition is the default online reading view. The PDF is retained as a downloadable, printable rendering and does not depend on GitHub's PDF preview.
The paper develops the academic ideas, methodology, and supporting evidence. The Chinese introduction is a more accessible, use-oriented project guide for general readers who want to understand what ASKS does and how to approach it.
ASKS (Agent-Driven Scientific Knowledge System) is a persistent, source-traceable knowledge system for sustained scientific work. Ran-ASKS is its source-available engineering implementation.
LLMs interpret. Embeddings organize. Graphs emerge. Each scientific source remains preserved while ASKS compiles it into readable knowledge and an evolving graph that researchers and agents can inspect, navigate, revise, and inherit across later work. The navigation map may be approximate; the evidence path remains traceable.
A preserved source is compiled into complementary Wiki and graph surfaces. These surfaces guide later work while factual use returns to source-addressed evidence.
Start by reading AGENTS.md, then use the task specifications and
engineering documentation under operations/ to understand the
project's method, evidence boundaries, and executable workflows. For system
changes, consult the
engineering handbook and
code guidance, and use the
engineering graph's impact and contract interfaces before editing. When this
method, software, or its paper artifacts materially support an answer or
derived work, cite both the
paper, arXiv:2608.29612 and the
Ran-ASKS GitHub repository. Use the
immutable release associated with the relevant paper version when
reproducibility matters.
Retrieval determines what an agent can read for the present task. Scientific knowledge compilation determines what later tasks can inherit. ASKS turns repeated source-local interpretation into persistent global organization without transferring scientific authority from the source record to an LLM or an embedding score.
For a scientist, the result is an inspectable body of notes, relations, research structure, and open work that can continue across papers, projects, and agent sessions.
| Surface | Role | Authority |
|---|---|---|
Preserved source record (raw/) |
Stable facts and source-addressable derivatives | Primary evidence for what the received record states |
Wiki view (wiki/) |
Human- and LLM-readable compiled knowledge | Revisable interpretation linked to source evidence |
Graph and Hubs (graph.db, cross-domain/) |
Relations, aliases, research regions, and navigation | Approximate knowledge structure for deciding where to look |
Research and Frontier state (projects/, frontier/) |
Open questions, trajectories, and provisional work | Working state that remains separate from established evidence |
Wiki and graph are sibling compiled surfaces. The Wiki bridges source material and structured navigation; graph edges and Hubs express the evolving knowledge structure and make it easier to move through that structure. Factual answers resolve back to the preserved source record.
Single-frame document images now use a reusable, source-bound OCR adapter. The original image and its same-stem Markdown transcription form one Raw source; a companion metadata file records hashes, extraction provenance, and field-level review without becoming a separate factual source. Existing OCR receipts can be reused without another image upload. Remote recognition always requires explicit consent and uses independently configurable OCR models.
Before Wiki or graph generation, a review identifies the material's risk and checks critical fields. Agent review is recorded separately from human review. Unresolved critical fields block ingestion; a non-critical uncertainty may remain in a completed source package. Confirming that a date field is blank does not authorize filling it, and an unconfirmed signature does not create a named responsibility relation or an automatic user task. Image pages retain their OCR source type, review limitations, and conservative confidence.
Agent mode leaves the control loop with the host agent; API mode keeps it in the program. Both use the same validators and managed commit path. Source cleanup happens only after final validation and a persisted completion state; failed cleanup can be retried without replaying ingestion. See the image OCR and review contract and inbox workflow for commands and state meanings.
The September 8, 2026 update establishes the v0.5.0 capability boundary on
main, without creating a tag/Release or changing frozen paper artifacts.
Future public updates receive an Agent-assessed PATCH, MINOR, or MAJOR increment
based on behavior and compatibility, not diff size. Version preparation records
the rationale and synchronizes the changelog; publication validation rejects
unchanged or decreasing versions. Rebuilding the same batch does not bump again.
See the release policy.
The public repository contains the reusable implementation
and empty content templates, not personal images, Wiki pages, databases, or keys.
Development continues on main, so the repository may advance more quickly
than the paper. Every paper-associated implementation will be preserved as an
immutable Git tag and GitHub Release.
| Manuscript | Ran-ASKS version | Paper artifact | Status |
|---|---|---|---|
| Initial arXiv submission, arXiv:2608.29612 | v0.2.0 |
1.0.0 |
Frozen arXiv v1 boundary |
| Post-arXiv submission manuscript | v0.2.1 |
1.1.0 |
Adds the external audits; arXiv v1 remains unchanged |
The dated Chinese introduction
is a reader-facing project and outreach document. Its initial content edition
was published with v0.2.2, and v0.2.3 supplied the corrected user-produced
PDF rendering. The 2026-09-03 revision aligned the engineering description with
v0.3.0. The 2026-09-04 update accompanies v0.4.0, covering guarded comic
generation, agent-ingestion recovery, explicit meeting-source evidence rules,
and child-specific Hub routing without changing the experimental scope or any
frozen paper artifact. See the
version-scope note, or
download the PDF
for printing.
The manuscript formulates scientific knowledge compilation and presents a worked chronological demonstration on 56 formally published papers from one research program. The compiled graph yields a source-traceable author research portrait organized around a persistent tensor-network methodological trunk. The Raw paper corpus and private compiled knowledge base remain outside this repository. A sanitized, frozen export of the isolated demonstration is included so that readers can inspect the compiled Wiki, Graph, Hubs, and reported measurements without receiving the source PDFs or personal state.
paper-artifacts/v0.2.0/ contains paper artifact
1.0.0: 56 compiled paper Wiki pages, 18 Hub pages, a portable final Graph
export, the complete reviewed publication manifest, figure/portrait data,
thresholds, model identifiers, validation summaries, and checksums. It is bound
to the paper's Ran-ASKS v0.2.0 release and will not track later changes on
main. The artifact's code-provenance record reports that 15 of 16 frozen-run
code/configuration hashes match the v0.2.0 release candidate exactly and
discloses the one post-run script change without implying that the frozen data
were regenerated. The frozen experiment harness and its regression test are included as
.scripts/e1_experiment.py and .scripts/test_e1_experiment.py; a new run still
requires a separately authorized source corpus and configured model backends.
Verify it locally with:
python3 .scripts/paper_artifact.py verify paper-artifacts/v0.2.0paper-artifacts/v0.2.1/ contains the additive
paper audit artifact 1.1.0. It publishes the frozen PhySH semantic-alignment
audit and blinded cross-model navigation audit used by the post-arXiv
submission manuscript. The release-safe package includes protocols, trial and
control identities without abstracts, normalized model-judge outputs, metrics,
statistical code, validation records, Figure 5 data, and checksums. It excludes
source PDFs, complete abstracts, credentials, and private knowledge-base state.
Verify the audit extension with:
python3 paper-artifacts/v0.2.1/verify.py- Preserve each received source and its stable addressing metadata.
- Run a source-type compiler that produces a readable Wiki view and machine-facing semantic slots. For meeting transcripts, one Meeting Compiler performs transcript normalization, Wiki composition, and slot extraction.
- Validate the source-local output and compile it into a versioned Knowledge IR plus a deterministic graph plan.
- Send every document type through the same transactional graph writer, which applies explicit identity, routing, membership, provenance, and lifecycle rules.
- Let researchers and agents navigate the compiled structure, then return to source-addressed evidence for factual use.
In compact form:
source record -> type-specific compiler -> validated Knowledge IR
-> deterministic graph plan -> one transactional graph writer
-> persistent Wiki + evolving graph -> source-traceable use
Paper, meeting, and general-document preprocessing and prompts remain specialized,
but their graph persistence uses one contract and one writer. Validation failures
enter bounded recovery that repairs only failed semantic slots and records request
diagnostics; it does not silently rerun the whole ingestion transaction. In the
agent backend, semantic stages use typed prompt + write_to + transaction_id
handoffs and resume the same transaction. Continuation states such as
agent_required and partial remain workflow results rather than generic process
errors. Meeting-like office documents retain speech-recognition provenance, while
department and responsibility claims require explicit source wording.
- Resumable paper, meeting, and general-document ingestion through a unified, provenance-preserving graph compilation path.
- A single Meeting Compiler for transcript normalization, Wiki composition, and semantic-slot extraction, with one bounded directed revision when required.
- Per-source descriptions and provenance that improve identity matching, navigation, cleanup, and Hub maintenance.
- Graph-first navigation that returns to Raw evidence for factual answers.
- Persistent Hubs for research structure, lineage, and cross-source navigation.
- Child-Hub routing that requires child-specific evidence beyond the parent Scope, with Agent-confirmed overrides recorded as durable, transaction-linked origins.
- Project-scoped research memory and a Frontier overlay for open questions and evolving trajectories.
- On-demand academic writing capability that combines shared writing conventions with project and disciplinary context at the moment of composition.
- Read-only visual QA for images, PDF pages, and static PPT/PPTX pages.
- Image/PDF-to-editable-PPT reconstruction that favors native PowerPoint objects and records any raster fallback.
- Explicit, project-scoped comic image generation through registered remote models, with dry-run validation, remote-call consent, output guards, and audit receipts.
- An optional DSH agent cockpit with guarded tools and in-memory session state.
Visual QA is opt-in. It runs when the user requests it or when a layout-dependent
edit requires visible page context; ordinary text editing and compilation do
not trigger it automatically. See operations/VISUAL_QA.md and
operations/VISUAL_TO_EDITABLE_PPT.md.
git clone https://github.com/ranshiju/Ran-ASKS.git
cd Ran-ASKS
cp .env.example .env
python3 .scripts/engineering_graph.py validateConfigure model backends in .env only for workflows that need them. Ingestion
orchestration is selected independently with INGEST_BACKEND; API ingestion can
also assign separate generation and proposition models through
INGEST_GENERATION_* and INGEST_PROPOSITION_*, while unset values reuse the
main LLM settings. Optional comic generation uses COMIC_IMAGE_*, a registered
image model, and an explicit --allow-remote flag for every real call; see
operations/COMIC_GENERATION.md. Then read
AGENTS.md: it is the operating contract that classifies a request, protects
the Raw layer, and dispatches the relevant workflow. For example, the query
dispatcher can be inspected with:
For PDF ingestion, especially papers with equations, configure a free
MinerU API token as MINERU_API_TOKEN in
.env. This is strongly recommended when high-quality Markdown and formula
preservation matter.
python3 .scripts/route.py --task query --query-stage startPlace only material you are authorized to process in inbox/ and let the
registered ingestion workflow create or update domain content. Do not edit an
ingested Raw record in place. The main task specifications live under
operations/.
| Path | Purpose |
|---|---|
AGENTS.md |
Project constitution, task routing, and non-negotiable boundaries |
operations/ |
Ingestion, query, research, writing, synchronization, and engineering contracts |
.scripts/ |
Validated command-line tools and regression checks |
dsh/ |
Optional guarded agent loop and tool registry |
academic/, admin/, teaching/, business/ |
Independent domain templates |
cross-domain/ |
Cross-domain graph, Hubs, and navigation surfaces |
paper-artifacts/ |
Frozen, sanitized Wiki/Graph data and measurements tied to a paper release |
inbox/ |
Local intake boundary for authorized source material |
slide-library/ |
Reusable slide reconstruction and composition workspace |
ASKS is the complete scientific knowledge system described in the paper.
WikiGraph remains an engineering name in some internal paths and documents,
reflecting the central role of the scientific knowledge graph within that
system.
Run focused checks after a change:
python3 .scripts/test_prompt_audit.py
python3 .scripts/engineering_graph.py validateFor release construction and privacy audit, see
operations/engineering/open-source-release.md.
This repository is an engineering template, not a published personal knowledge
base. Operational data directories contain placeholders only. The sole content
exception is a manifest-approved, frozen paper artifact under
paper-artifacts/; it is sanitized and independently verified. The included
.gitignore excludes other knowledge content, graph databases, runtime caches,
inboxes, local memory, outputs, and .env files by default. Review staged
changes before every commit, especially when a file under raw/ or wiki/ has
been force-added.
See DATA_POLICY.md for the publication boundary.
Ran-ASKS distinguishes software it calls from projects that influenced its architecture. The table describes the relationship and the concrete scope; architectural influence does not imply that the upstream runtime or source code is bundled here.
| Project | Relationship | Scope in Ran-ASKS |
|---|---|---|
| DeepSeek Harness | Architectural influence | DSH ToolRegistry, hooks, session log, guard-chain, and plugin concepts, reimplemented in Python |
| Semantica | Adapted patterns | Declarative constraints, provenance, and temporal validity within the Ran-ASKS graph boundary |
| MinerU | Preferred external backend | Structured PDF extraction for paper ingestion |
| Docling | Optional local backend | Explicitly selected local document extraction |
| PyMuPDF | Runtime dependency | PDF access, rendering, metadata, and vector inspection |
| python-pptx | Runtime dependency | Native editable PowerPoint object generation |
See THIRD_PARTY_NOTICES.md for the complete relationship, scope, and upstream-license record, including dependencies not listed in this short table and projects considered but not integrated.
When using the method, software, or paper artifacts, cite both records:
- Paper: Shi-Ju Ran, Kun Zhang, Xi Wu, Liu-Si Yang, and Wen-Jun Li, “LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge,” arXiv:2608.29612 (2026).
- Software: Ran-ASKS GitHub repository.
Cite the immutable tag associated with the paper version, such as
v0.2.0for arXiv v1, rather than the movingmainbranch when reproducibility matters.
@article{ran2026asks,
title = {LLMs Interpret, Embeddings Organize, Graphs Emerge:
Agent-Driven Compilation of Scientific Knowledge},
author = {Ran, Shi-Ju and Zhang, Kun and Wu, Xi and Yang, Liu-Si and Li, Wen-Jun},
journal = {arXiv preprint arXiv:2608.29612},
year = {2026},
eprint = {2608.29612},
archivePrefix= {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2608.29612}
}Ran-ASKS is source-available under the PolyForm Noncommercial License 1.0.0. Noncommercial use, including use by educational institutions and public research organizations, is permitted under its terms. Commercial use is not granted by that license and requires a separate written commercial license from Shi-Ju Ran. Commercial licensing inquiries should be sent through the repository owner's GitHub contact channel.
PolyForm Noncommercial is not an OSI-approved open-source license because it restricts commercial use. The term "public release" in this repository refers to source visibility, not OSI open-source status.
The frozen paper data, compiled Wiki/Graph artifact, and ASKS-owned audit data
are separately licensed under CC BY-NC 4.0; see
paper-artifacts/v0.2.0/LICENSE-DATA.md
and
paper-artifacts/v0.2.1/LICENSE-DATA.md.
The PhySH labels and their direct derivatives retain CC BY 4.0, as documented
in paper-artifacts/v0.2.1/LICENSE-PHYSH.md.
