# OSCODE Workspace

Open **작업실** (Workspace) in the web CLI. Controls currently use Korean labels. The following ten features operate on actual records or explicitly supplied data; finished responses, passing checks and pixel differences are distinct.

1. **Recorded flow → test draft.** In the work browser, choose **동작 테스트 초안 기록 시작**, consent to recording ordinary input values, perform actions, then choose **동작 테스트 초안 받기**. Download a Playwright `@playwright/test` draft; review selectors and add business assertions before running it. Password fields are omitted; ordinary fields may contain sensitive data. At most 100 actions are recorded. The default assertion checks the URL only. The `test-export` API accepts `checks` for generated URL, visibility and text assertions. Drafts are never automatically executed.
2. **Completion checks.** Specify 1–12 JSON checks using `visible` (selector), `text` (selector/value), `url` (exact value), or `noErrors`. Run them against the manually controlled work browser. Error checks cover collected browser errors only; reset diagnostic history before a fresh check. Unspecified behavior is not considered verified.
3. **Module canvas.** Connect file-reading, AI analysis, generation, actual completion checks and human review blocks. Reorder with up/down buttons. File blocks take a JSON array of 1–8 project paths; check blocks take the checks above. Selected files retain the existing 24KB/file and 48KB total limits and project/secret boundaries. File content is untrusted context, not instructions. Each step requires approval. Failed check blocks prevent advancement and can be checked again after fixing the screen. Definitions persist; execution state does not survive restart.
4. **Change impact map.** Enter changed paths, one per line. Relative imports provide transitive module/page/test candidates and dependency evidence. Up to 300 source files and 4MB are inspected. Aliases, computed dependencies and external packages are not fully resolved; unresolved paths are reported. This is not a complete dependency proof.
5. **Two implementation workspaces.** Approve two prompts to run them in separate file copies and web CLI instances. Process tool approvals in each window. Compare task records, file hashes, recorded costs and two development-screen URLs. This does not create Git branches or merge changes into the original. Maximum three pairs, 1,000 files/64MiB per copy and 2MiB per file. Dependencies, build outputs, known credential settings, environment files and symlinks are omitted; review omissions. Embedded secrets cannot be automatically identified. Dependency installation/server startup need separate approvals. Each window retains the 30-second disconnected-idle cancellation policy. Copies remain at the displayed temporary paths after the parent server closes. External side effects are not represented by file comparisons.
6. **Grouped failures.** Group actual failed tool outputs, console errors, page exceptions, request failures and HTTP errors. Inspect counts, original evidence IDs and diagnostic suggestions. Preparing a rework request creates a draft only. Suggestions do not establish root cause. Browser diagnostics retain the latest 100 records in memory.
7. **Verification evidence package.** Select 1–10 tasks and approve a ZIP download containing `report.json`, notes, statuses, tool-change paths, approvals, estimated costs and recorded checks. Optionally include up to five snapshots and the baseline. Source export includes up to 30 tool checkpoints with pre-change/current content capped at 65,536 characters each and truncation flags. Current content is read at export time and may include subsequent changes. Shell changes may be absent. Maximum ZIP size: 16MiB. Review sensitive data before sharing; no automatic external upload occurs. Browser checks are not automatically attested to belong to every selected task.
8. **Reuse finished work.** Propose a recipe from a finished task, review/edit it, save it, then separately approve execution. Existing approvals or tool calls are not replayed. Response completion is not a quality guarantee.
9. **Event assistant.** Save an exact file path, a one-off ISO timestamp, or a current-session `mcp_…` tool completion event with a suggested prompt. Explicitly enable observation. Events only create proposals; refresh and approve or dismiss them. MCP completion observation is not remote push subscription or polling. Up to ten definitions persist in `.oscode/studio-events.json`; watchers and up to fifty proposals are in memory. Manually re-enable after restart. Past one-off timestamps produce a new proposal when enabled again. No recurring scheduling or collection while the server is offline. File proposals are limited to one per second.
10. **Semiconductor model comparison and saved failures.** In **반도체 모델 비교**, choose **검사 실패 사례 모으기** to open the synthetic inspection environment in a new tab. The **가상 생산라인** panel offers the same link; the existing Model Lab remains available for general image classification, detection and segmentation. Run a synthetic LOT, select a completed case with false positives or negatives, enter a suite name, then choose **선택한 실패 저장** in **검사 회귀 작업실**. This freezes the exact requested PNG, truth, baseline predictions and the original run's score/IoU thresholds. Optionally edit the rectangle JSON and check **정답을 검수했습니다**. Save/reopen the suite with **회귀 세트 저장 / 회귀 세트 열기**; limits are 48 cases, 12 MiB of PNG files and 18 MiB of JSON. Data stays in browser memory, so save before refreshing. Configure an existing MLflow detection model, then choose **새 모델로 회귀 검사** and confirm transmission of the frozen PNGs; truth and baseline predictions are not sent. Review remaining, reduced or worsened errors and save the result with **검사 결과 저장**. Passing requires every case to complete, no per-case increase in FP or FN, and the total FP/FN limits (both default to zero). Missing, failed or cancelled cases leave the run incomplete. This evaluates the selected failures; it does not establish overall or production performance or approve deployment. Human review and model ID/version are user records, not external certification or verified server identity. See the [inspection regression guide](inspection-simulator.md#오류-사례를-회귀-검사로-재사용). This is a local development feature; check the installed/published package version for availability.

### Advanced: compare supplied JSON

The same **반도체 모델 비교** panel retains the existing JSON fields below the saved-failure link. **동일 데이터를 두 모델로 전송 승인** sends one identical JSON payload sequentially to two local HTTP(S) endpoints, including MLflow `/invocations`, and measures actual response times. **가져온 예측 평가** evaluates two supplied prediction arrays without calling a server. This is separate from the frozen PNG regression suites.

Responses must contain `predictions` records with unique inspection IDs. Classification records use `{id,label,score}` and truth uses `{id,label}`. False positives/negatives are evaluated against the configured positive label. Detection records use `{id,objects:[{label,box:[x,y,width,height],score}]}`; truth objects omit scores. Match detections one-to-one with confidence and IoU thresholds in the same image coordinates. Imported results are not labeled as live inference. Without ground truth, accuracy and false-positive/negative metrics remain unknown. Maximum 500 inspection IDs, 100 objects per inspection, 512KB request, 1MiB response, 15 seconds/model. Optional authentication uses an endpoint `tokenEnv` environment-variable name, never a saved token value. Matrix/custom responses require adapters. A single run does not establish production latency or generalization; no automatic MES/equipment connection is performed.

Example completion checks:

```json
[{"kind":"visible","selector":"#app"},{"kind":"text","selector":"#result","value":"ready"},{"kind":"noErrors"}]
```

Example classification response:

```json
{"predictions":[{"id":"LOT-01-die-3","label":"defect","score":0.91}]}
```

See [reliable workflows](reliable-workflows.en.md) for recovery, analysis reuse, preference policies, model gates, file promotion and offline replay.
