Yamma · Claude development plan

Build carefully. Test against the record.

Our Claude plan covers coding, test design, and a separate evaluation of Korean structured extraction and factual summaries.

Development plan · Evaluation results will be published after execution.

Claude in development.

Our development workflow includes Claude Code guidance for coding and test design. The next proposed step is an isolated evaluation using authored development cases.

Development support

Code, tests, and documentation.

Use developer review and repository checks before accepting software changes.

Planned evaluation

Structured symptom extraction.

Compare candidate fields with the supplied note. Keep occurrence dates, recording dates, and uncertain details distinct.

Planned evaluation

Factual summary drafts.

Check each statement against the source record, including missing facts, negation, and corrections.

An evaluation with a defined boundary.

Prepare 100 authored cases across five groups: symptom combinations, dates, missing or uncertain information, negation or conflicts, and malformed or out-of-scope input. These are development scenarios, not health records from actual users.

Planned run
Run each case three times for extraction and summary drafting: 600 planned requests per model and prompt configuration. Keep prompt-development examples separate.
Review flow
Authored case → candidate output → schema and source comparison → human review → evaluation report.
What we measure
Schema validity, preservation of known and unknown fields, unsupported additions, factual omissions, error handling, latency, and cost.
Acceptance boundary
Unsupported claims, critical errors, or unsafe failure paths block acceptance. Reviewers compare output with reference facts and retain earlier results after corrections.

No diagnosis, triage decision, medication change, or unsupported fact belongs in a test output. API responses stay test artifacts and do not automatically become saved records.

PDF · English · 2 pages · Updated 10 October 2026

The product data path is a separate decision.

The implemented manual-entry path does not require a Claude call. The proposed evaluation does not change the current production direction.

Current manual flow
Entry on the Android device → user review and confirmation → encrypted local storage → diary and intensity timeline.
Planned production direction
Self-hosted open models in South Korea. A future model or operating-architecture decision would require its own review.
Separate evaluation
Use authored cases in an isolated workflow, without real health records, identifiers, or application secrets. Measure actual usage and spending.

Publish what we learn.

The proposed report will include methods, case definitions, counts, failures, limitations, latency, and cost. The plan contains no completed benchmark or clinical-validation result.

Official documentation informs the method: structured outputs, evaluation design, and pricing.