NAP2027 kā dati: shēma strategija-0.1, lv-nap2027, datu līgums B06
475 numurētie punkti (mērķi, 131 indikators, 124 uzdevumi ar VPK ID, telpiskās attīstības virzieni), 154 pamatojuma ieraksti, 34 zemsvītras piezīmes. Pārbaude pret avota PDF ar citu nolasītāju — izturēta. Katalogs planosanas-dokumenti.yaml, rīki parse/build/verify/validate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016679RwmHsuTFfxt26wP6rk
This commit is contained in:
12
CHANGELOG.md
Normal file
12
CHANGELOG.md
Normal file
@@ -0,0 +1,12 @@
|
|||||||
|
# Izmaiņas
|
||||||
|
|
||||||
|
## 0.1.0 — 2026-10-11
|
||||||
|
|
||||||
|
- Shēma `schemas/strategija-0.1.xsd` attīstības plānošanas dokumentiem: nodaļas, numurētie punkti ar veidu, indikatori,
|
||||||
|
uzdevumi ar institūcijām (VPK ID), finansējuma avoti, uzdevumu indikatoru saites, pamatojuma avoti, zemsvītras piezīmes.
|
||||||
|
- `lv-nap2027` — Latvijas Nacionālais attīstības plāns 2021.–2027. gadam: 475 punkti, 154 pamatojuma ieraksti, 34 piezīmes.
|
||||||
|
Pārbaudes atskaite `verification/lv-nap2027.verify.json`.
|
||||||
|
- Institūciju sasaiste `sources/nap2027/dalibnieki.yaml`; Valsts institūciju reģistrā (Valdibas-Deklaracija-as-Code,
|
||||||
|
2. datu versija) pievienotas NAP2027 institūcijas un grupas, kuru tur nebija.
|
||||||
|
- Katalogs `planosanas-dokumenti.yaml`, datu līgums `contracts/B06-attistibas-planosanas-dokumenti.odcs.yaml`,
|
||||||
|
rīki `tools/parse_nap.py`, `tools/build_nap2027.py`, `tools/verify_nap2027.py`, `tools/validate.py`.
|
||||||
91
README.md
91
README.md
@@ -1,3 +1,92 @@
|
|||||||
# Strategy-as-Code
|
# Strategy-as-Code
|
||||||
|
|
||||||
Attīstības plānošanas dokumenti kā dati (PPPA, Valsts Pirmkods). Melnraksts — saturs tiek pievienots.
|
Attīstības plānošanas dokumenti kā dati. Katrs apstiprināts dokuments — viena XML datne, kas atbilst shēmai
|
||||||
|
`schemas/strategija-0.1.xsd`: nodaļas, numurētie punkti, mērķi, indikatori ar bāzes un mērķa vērtībām, uzdevumi ar atbildīgajām
|
||||||
|
institūcijām, finansējuma avoti, pamatojuma avoti un zemsvītras piezīmes.
|
||||||
|
|
||||||
|
PPPA koncepcijas demonstrācija (Valsts PirmKods), nevis oficiāls izdevums. Juridiski saistošs ir apstiprinātais dokuments;
|
||||||
|
tā PDF un sha256 ir repozitorijā.
|
||||||
|
|
||||||
|
## Saturs
|
||||||
|
|
||||||
|
| Ceļš | Kas tas ir |
|
||||||
|
|---|---|
|
||||||
|
| `planosanas-dokumenti.yaml` | Dokumentu katalogs: identifikators, apstiprinājums, avots, datne |
|
||||||
|
| `schemas/strategija-0.1.xsd` | Shēma (vārdtelpa `urn:pppa:vpk:strategija:0.1`) |
|
||||||
|
| `contracts/B06-attistibas-planosanas-dokumenti.odcs.yaml` | Datu līgums (ODCS v3) — ko un kā ielādē VDZP (lakehouse.pppa.lv) |
|
||||||
|
| `data/<dokuments>/<id>.xml` | Dokuments kā dati |
|
||||||
|
| `sources/<dokuments>/` | Avota PDF un institūciju apzīmējumu sasaiste ar VPK ID |
|
||||||
|
| `verification/<id>.verify.json` | Datnes pārbaude pret avota PDF |
|
||||||
|
| `tools/` | Nolasīšana no PDF, XML veidošana, pārbaude |
|
||||||
|
|
||||||
|
## Dokumenti
|
||||||
|
|
||||||
|
| ID | Dokuments | Apstiprināts | Stāvoklis |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `lv-nap2027` | Latvijas Nacionālais attīstības plāns 2021.–2027. gadam | Saeima, 02.07.2020., lēmums Nr. 418/Lm13 | pārveidots |
|
||||||
|
|
||||||
|
### lv-nap2027
|
||||||
|
|
||||||
|
- 475 numurētie punkti [1]–[475]: 165 teksta punkti, 4 stratēģiskie mērķi, 6 prioritāšu mērķi, 25 rīcības virzienu mērķi,
|
||||||
|
131 indikators, 124 uzdevumi, 20 telpiskās attīstības virzieni.
|
||||||
|
- 6 prioritātes, 18 rīcības virzieni (numurēti visā dokumentā: `pr1.rv1` … `pr6.rv18`), rīcības virzienā „Kvalitatīva,
|
||||||
|
pieejama, iekļaujoša izglītība” — 4 tematiskas apakšnodaļas (`pr2.rv6.t1`–`t4`). Katram rīcības virzienam — indikatīvi
|
||||||
|
pieejamais finanšu apjoms, kā drukāts.
|
||||||
|
- Pielikums: 152 numurēti pamatojuma punkti un 2 nenumurētas rindkopas, ar problēmu, pie kuras grupēti, un tīmekļa adresēm.
|
||||||
|
- 34 zemsvītras piezīmes ar atsaucēm no punktiem; 66 saīsinājumi.
|
||||||
|
|
||||||
|
## Identifikatori
|
||||||
|
|
||||||
|
| Elements | Veidne | Piemērs |
|
||||||
|
|---|---|---|
|
||||||
|
| Dokuments | `lv-<saīsinājums>` | `lv-nap2027` |
|
||||||
|
| Nodaļa | `<dok>.prN`, `<dok>.prN.rvM`, `<dok>.prN.rvM.tK`, `<dok>.<veids>` | `lv-nap2027.pr4.rv11`, `lv-nap2027.spatial` |
|
||||||
|
| Punkts | `<dok>.iNNN` — drukātais numurs | `lv-nap2027.i068` |
|
||||||
|
| Pamatojums | `<dok>.aNNN`; nenumurēta rindkopa `<dok>.annex.uN` | `lv-nap2027.a078` |
|
||||||
|
| Institūcija | VPK ID no Valsts institūciju reģistra | `21-0000` |
|
||||||
|
|
||||||
|
Identifikatori nemainās starp datu versijām un netiek izmantoti atkārtoti.
|
||||||
|
|
||||||
|
## Institūcijas
|
||||||
|
|
||||||
|
Uzdevuma atbildīgās un līdzatbildīgās institūcijas datnē ir ar drukāto apzīmējumu (`label`) un VPK ID (`org`) no
|
||||||
|
[Valsts institūciju reģistra](https://processgit.org/Valsts-Pirmkods/Valdibas-Deklaracija-as-Code/src/branch/main/data/organizacijas.xml).
|
||||||
|
Sasaiste ir `sources/nap2027/dalibnieki.yaml`; elementā `Actors` katram apzīmējumam norādīts sasaistes veids:
|
||||||
|
|
||||||
|
| `match` | Nozīme | NAP2027 |
|
||||||
|
|---|---|---|
|
||||||
|
| `direct` | Tā pati institūcija | 25 apzīmējumi |
|
||||||
|
| `renamed` | Pārdēvēta; reģistrā iepriekšējais nosaukums | VARAM (no 2024-07-01 Viedās administrācijas un reģionālās attīstības ministrija) |
|
||||||
|
| `historical` | Likvidēta vai apvienota; reģistrā `validTo` un `successor` | PKC, PKC (DLC) → Valsts kanceleja; FKTK → Latvijas Banka |
|
||||||
|
| `group` | Institūciju grupa | visas ministrijas, pašvaldības, NVO, plānošanas reģioni u. c. |
|
||||||
|
| `assumed` | Apzīmējums nav dokumenta saīsinājumu sarakstā; nozīme pieņemta | LTKR (→ LTRK), LABS (→ LBAS), LIZDA, LLPA |
|
||||||
|
| `none` | Nav valsts institūcija | Nasdaq |
|
||||||
|
|
||||||
|
Uzdevumi saista institūciju, kas norādīta apstiprinātajā dokumentā. Ja funkcija vēlāk pārcelta citai institūcijai
|
||||||
|
(piemēram, vides aizsardzības politika no VARAM uz Klimata un enerģētikas ministriju 2024. gadā), datnē paliek dokumentā
|
||||||
|
norādītā institūcija; pārmaiņa ir reģistrā.
|
||||||
|
|
||||||
|
## Uzdevumu indikatori
|
||||||
|
|
||||||
|
Uzdevumu tabulas ailē „Indikators” norādītais nosaukums sasaistīts ar dokumenta mērķa indikatoru (`IndicatorRef`), ja
|
||||||
|
nosaukums sakrīt (`exact`), ir saīsināts (`shortened`) vai gandrīz sakrīt (`similar`). 37 nosaukumiem dokumentā nav mērķa
|
||||||
|
indikatora — tie ir uzdevuma līmeņa indikatori tikai ar nosaukumu.
|
||||||
|
|
||||||
|
## Kā veidota un pārbaudīta datne
|
||||||
|
|
||||||
|
```
|
||||||
|
python3 tools/parse_nap.py sources/nap2027/NAP2027.pdf build/nap2027.json # nolasīšana (pdfplumber)
|
||||||
|
python3 tools/build_nap2027.py build/nap2027.json data/nap2027/lv-nap2027.xml # XML, VPK ID, indikatoru saites
|
||||||
|
python3 tools/verify_nap2027.py data/nap2027/lv-nap2027.xml sources/nap2027/NAP2027.pdf verification/lv-nap2027.verify.json
|
||||||
|
python3 tools/validate.py
|
||||||
|
```
|
||||||
|
|
||||||
|
Pārbaude izmanto citu PDF nolasītāju (poppler `pdftotext`) nekā datnes veidošana: katrs drukātais numurs datnē ir tieši
|
||||||
|
vienreiz; teksta punktu un pamatojuma punktu burti ir avota tekstā (dažiem — 2–4 daļās, jo avota lasīšanas secībā starp
|
||||||
|
tām ir zemsvītras piezīme vai tabulas aile); katrs tabulas rindas vārds (9684) ir avotā starp punkta numuru un nākamo numuru;
|
||||||
|
katram uzdevumam ir atbildīgā institūcija ar VPK ID; katram indikatoram — vērtības.
|
||||||
|
|
||||||
|
## Datu kopa VDZP
|
||||||
|
|
||||||
|
Datu kopa B06 VDZP (lakehouse.pppa.lv) ielādē šo repozitoriju katru dienu (ProcessGit savienotājs, datu līgums
|
||||||
|
`contracts/B06-…`).
|
||||||
|
|||||||
139
contracts/B06-attistibas-planosanas-dokumenti.odcs.yaml
Normal file
139
contracts/B06-attistibas-planosanas-dokumenti.odcs.yaml
Normal file
@@ -0,0 +1,139 @@
|
|||||||
|
# ODCS v3 datu līgums: B06 Attīstības plānošanas dokumenti (stratēģija kā kods)
|
||||||
|
# Shēma: schemas/strategija-0.1.xsd. Šī līguma schema daļa apraksta to, ko VDZP ielādē; pilnā struktūra ir XSD.
|
||||||
|
apiVersion: v3.0.2
|
||||||
|
kind: DataContract
|
||||||
|
id: urn:pppa:cac:contract:B06
|
||||||
|
name: Attīstības plānošanas dokumenti (stratēģija kā kods)
|
||||||
|
version: 0.1.0
|
||||||
|
status: active
|
||||||
|
domain: Valsts PirmKods
|
||||||
|
dataProduct: Attīstības plānošanas dokumenti
|
||||||
|
tenant: PPP Asociācija (PPPA)
|
||||||
|
description:
|
||||||
|
purpose: >-
|
||||||
|
Apstiprināti attīstības plānošanas dokumenti kā dati — pirmais ir Nacionālais attīstības plāns 2021.–2027. gadam (lv-nap2027):
|
||||||
|
prioritātes, rīcības virzieni, mērķi, indikatori ar bāzes un mērķa vērtībām, uzdevumi ar atbildīgajām un līdzatbildīgajām
|
||||||
|
institūcijām (VPK ID), finansējuma avoti, indikatīvais finansējums, telpiskās attīstības virzieni un pamatojuma avoti.
|
||||||
|
usage: Mašīnlasāma datu apmaiņa; ielāde VDZP katru dienu; AI aģentu vaicājumi caur MCP.
|
||||||
|
limitations: >-
|
||||||
|
Koncepcijas demonstrācija, nav oficiāls izdevums. Juridiski saistošs ir apstiprinātais dokuments (avota PDF ar sha256).
|
||||||
|
Teksts un vērtības — kā drukāti; datnes atbilstība avotam pārbaudīta ar citu PDF nolasītāju (verification/<id>.verify.json).
|
||||||
|
Institūciju sasaiste ar VPK ID: drukātais apzīmējums saglabāts; pārdēvētām un likvidētām institūcijām — reģistra ID ar pēcteci;
|
||||||
|
pieņēmumi atzīmēti (match = assumed).
|
||||||
|
servers:
|
||||||
|
- server: processgit
|
||||||
|
type: custom
|
||||||
|
format: xml
|
||||||
|
description: ProcessGit repozitorijs (Gitea API v1)
|
||||||
|
customProperties:
|
||||||
|
- {property: url, value: "https://processgit.org/Valsts-Pirmkods/Strategy-as-Code"}
|
||||||
|
- {property: ref, value: main}
|
||||||
|
- {property: catalogue, value: planosanas-dokumenti.yaml}
|
||||||
|
- {property: path, value: "data/*/lv-*.xml"}
|
||||||
|
- {property: schema, value: "https://processgit.org/Valsts-Pirmkods/Strategy-as-Code/src/branch/main/schemas/strategija-0.1.xsd"}
|
||||||
|
schema:
|
||||||
|
- name: PlanningDocument
|
||||||
|
logicalType: object
|
||||||
|
physicalType: xml-element
|
||||||
|
physicalName: "/PlanningDocument"
|
||||||
|
customProperties: [{property: lakehouseTable, value: cac_bronze.sac_document}]
|
||||||
|
properties:
|
||||||
|
- {name: doc_id, logicalType: string, required: true, primaryKey: true, physicalName: "@id"}
|
||||||
|
- {name: title, logicalType: string, required: true, physicalName: Metadata/Title}
|
||||||
|
- {name: document_type, logicalType: string, required: true, physicalName: Metadata/DocumentType}
|
||||||
|
- {name: period_from, logicalType: integer, required: true}
|
||||||
|
- {name: period_to, logicalType: integer, required: true}
|
||||||
|
- {name: approved, logicalType: date, required: true, physicalName: Metadata/Approval/Date}
|
||||||
|
- {name: source_sha256, logicalType: string, required: true, physicalName: Metadata/Source/SHA256}
|
||||||
|
- name: Section
|
||||||
|
logicalType: object
|
||||||
|
physicalType: xml-element
|
||||||
|
physicalName: "//Section"
|
||||||
|
customProperties: [{property: lakehouseTable, value: cac_bronze.sac_section}]
|
||||||
|
properties:
|
||||||
|
- {name: section_id, logicalType: string, required: true, primaryKey: true, physicalName: "@id", description: "lv-nap2027.pr1, lv-nap2027.pr1.rv1"}
|
||||||
|
- {name: kind, logicalType: string, required: true, quality: [{type: library, rule: validValues, validValues: [introduction, vision, framework, strategicGoals, priority, actionLine, theme, spatial, implementation]}]}
|
||||||
|
- {name: parent_id, logicalType: string, required: false}
|
||||||
|
- {name: title, logicalType: string, required: true}
|
||||||
|
- {name: funding_million_eur, logicalType: number, required: false, physicalName: IndicativeFunding/@millionEur}
|
||||||
|
- name: Item
|
||||||
|
logicalType: object
|
||||||
|
physicalType: xml-element
|
||||||
|
physicalName: "//Item"
|
||||||
|
customProperties: [{property: lakehouseTable, value: cac_bronze.sac_item}]
|
||||||
|
properties:
|
||||||
|
- {name: item_id, logicalType: string, required: true, primaryKey: true, physicalName: "@id", description: "lv-nap2027.i068"}
|
||||||
|
- {name: n, logicalType: integer, required: true, description: Drukātais punkta numurs}
|
||||||
|
- {name: kind, logicalType: string, required: true, quality: [{type: library, rule: validValues, validValues: [text, strategicGoal, priorityGoal, actionLineGoal, indicator, task, spatialDirection]}]}
|
||||||
|
- {name: section_id, logicalType: string, required: true}
|
||||||
|
- {name: text, logicalType: string, required: true, description: "Text; indikatoram — Name; uzdevumam — Task/Text"}
|
||||||
|
- {name: unit, logicalType: string, required: false}
|
||||||
|
- {name: base_year, logicalType: string, required: false, description: kā drukāts}
|
||||||
|
- {name: base_value, logicalType: string, required: false, description: kā drukāts}
|
||||||
|
- {name: target_2024, logicalType: string, required: false}
|
||||||
|
- {name: target_2027, logicalType: string, required: false}
|
||||||
|
- {name: data_source, logicalType: string, required: false}
|
||||||
|
- name: TaskActor
|
||||||
|
logicalType: object
|
||||||
|
physicalType: xml-element
|
||||||
|
physicalName: "//Task/Responsible/Actor | //Task/CoResponsible/Actor"
|
||||||
|
customProperties: [{property: lakehouseTable, value: cac_bronze.sac_task_actor}]
|
||||||
|
properties:
|
||||||
|
- {name: item_id, logicalType: string, required: true}
|
||||||
|
- {name: role, logicalType: string, required: true, quality: [{type: library, rule: validValues, validValues: [responsible, co-responsible]}]}
|
||||||
|
- {name: label, logicalType: string, required: true, description: Drukātais apzīmējums}
|
||||||
|
- {name: vpk_id, logicalType: string, required: false, logicalTypeOptions: {pattern: "^[0-9]{2}-[0-9]{4}$"}}
|
||||||
|
- {name: match, logicalType: string, required: true, quality: [{type: library, rule: validValues, validValues: [direct, renamed, historical, group, assumed, none]}]}
|
||||||
|
- name: TaskIndicator
|
||||||
|
logicalType: object
|
||||||
|
physicalType: xml-element
|
||||||
|
physicalName: "//Task/TaskIndicator"
|
||||||
|
customProperties: [{property: lakehouseTable, value: cac_bronze.sac_task_indicator}]
|
||||||
|
properties:
|
||||||
|
- {name: item_id, logicalType: string, required: true}
|
||||||
|
- {name: name, logicalType: string, required: true}
|
||||||
|
- {name: indicator_id, logicalType: string, required: false, description: "Dokumenta mērķa indikators (IndicatorRef/@ref)"}
|
||||||
|
- {name: match, logicalType: string, required: false}
|
||||||
|
- name: TaskFunding
|
||||||
|
logicalType: object
|
||||||
|
physicalType: xml-element
|
||||||
|
physicalName: "//Task/Funding/Source"
|
||||||
|
customProperties: [{property: lakehouseTable, value: cac_bronze.sac_task_funding}]
|
||||||
|
properties:
|
||||||
|
- {name: item_id, logicalType: string, required: true}
|
||||||
|
- {name: code, logicalType: string, required: true}
|
||||||
|
- {name: label, logicalType: string, required: true}
|
||||||
|
- name: Evidence
|
||||||
|
logicalType: object
|
||||||
|
physicalType: xml-element
|
||||||
|
physicalName: "//Annex/Evidence"
|
||||||
|
customProperties: [{property: lakehouseTable, value: cac_bronze.sac_evidence}]
|
||||||
|
properties:
|
||||||
|
- {name: evidence_id, logicalType: string, required: true, primaryKey: true}
|
||||||
|
- {name: n, logicalType: integer, required: false}
|
||||||
|
- {name: section_id, logicalType: string, required: true}
|
||||||
|
- {name: problem, logicalType: string, required: false}
|
||||||
|
- {name: text, logicalType: string, required: true}
|
||||||
|
- {name: urls, logicalType: array, required: false}
|
||||||
|
quality:
|
||||||
|
- {name: xsd_valid, type: custom, engine: xsd, implementation: schemas/strategija-0.1.xsd,
|
||||||
|
description: "Datne atbilst XSD shēmai (arī atslēgas: punktu numuri unikāli, atsauces uz indikatoriem, institūcijām, finansējuma avotiem un piezīmēm atrisinās)", dimension: conformity}
|
||||||
|
- {name: filename_is_id, type: custom, engine: cac-lakehouse, implementation: "app.cac_strategy:filename_is_id", description: "Datnes nosaukums = dokumenta ID", dimension: consistency}
|
||||||
|
- {name: verified_against_source, type: custom, engine: cac-lakehouse, implementation: "app.cac_strategy:verified",
|
||||||
|
description: "Pārbaudes atskaite (verification/<id>.verify.json) ir par to pašu avota PDF un izturēta: visi numurētie punkti, teksts burtiski, tabulu vārdi", dimension: accuracy}
|
||||||
|
- {name: task_has_responsible_vpk_id, type: custom, engine: cac-lakehouse, implementation: "app.cac_strategy:task_resp",
|
||||||
|
description: "Katram uzdevumam ir atbildīgā institūcija ar VPK ID", dimension: completeness}
|
||||||
|
- {name: vpk_id_in_register, type: custom, engine: cac-lakehouse, implementation: "app.cac_strategy:vpk_in_register",
|
||||||
|
description: "Katrs dokumentā lietotais VPK ID ir Valsts institūciju reģistrā (B01)", dimension: consistency}
|
||||||
|
slaProperties:
|
||||||
|
- {property: frequency, value: 1, unit: d, element: "pull ingest katru dienu 06:30 Rīgas laikā"}
|
||||||
|
- {property: retention, value: "visas versijas (git vēsture)"}
|
||||||
|
team:
|
||||||
|
- {role: owner, name: "Dokumenta apstiprinātājs (NAP2027 — Saeima); izstrādātājs — Pārresoru koordinācijas centrs, tagad Valsts kanceleja"}
|
||||||
|
- {role: data steward, name: PPP Asociācija (PPPA)}
|
||||||
|
support:
|
||||||
|
- {channel: issues, url: "https://processgit.org/Valsts-Pirmkods/Strategy-as-Code/issues"}
|
||||||
|
customProperties:
|
||||||
|
- {property: sourceId, value: B06}
|
||||||
|
- {property: identifiers, value: "dokuments lv-nap2027; nodaļa lv-nap2027.pr1.rv1; punkts lv-nap2027.i068; pamatojums lv-nap2027.a001"}
|
||||||
|
- {property: validator, value: tools/validate.py}
|
||||||
6431
data/nap2027/lv-nap2027.xml
Normal file
6431
data/nap2027/lv-nap2027.xml
Normal file
File diff suppressed because it is too large
Load Diff
16
planosanas-dokumenti.yaml
Normal file
16
planosanas-dokumenti.yaml
Normal file
@@ -0,0 +1,16 @@
|
|||||||
|
# Attīstības plānošanas dokumentu katalogs (Strategy-as-Code).
|
||||||
|
# Viens ieraksts katram dokumentam; as_code.path — datne, ja dokuments pārveidots datos.
|
||||||
|
# Identifikators: lv-<saīsinājums> (mazie burti, bez atstarpēm), piemēram lv-nap2027.
|
||||||
|
documents:
|
||||||
|
- id: lv-nap2027
|
||||||
|
title: Latvijas Nacionālais attīstības plāns 2021.–2027. gadam
|
||||||
|
short_title: NAP2027
|
||||||
|
type: nacionālais attīstības plāns
|
||||||
|
period: {from: 2021, to: 2027}
|
||||||
|
approved: {body: Latvijas Republikas Saeima, act: Saeimas lēmums, number: 418/Lm13, date: 2020-07-02}
|
||||||
|
source: {url: "https://www.mk.gov.lv/lv/media/15162/download", file: sources/nap2027/NAP2027.pdf}
|
||||||
|
as_code:
|
||||||
|
status: transformed
|
||||||
|
path: data/nap2027/lv-nap2027.xml
|
||||||
|
schema: schemas/strategija-0.1.xsd
|
||||||
|
verification: verification/lv-nap2027.verify.json
|
||||||
336
schemas/strategija-0.1.xsd
Normal file
336
schemas/strategija-0.1.xsd
Normal file
@@ -0,0 +1,336 @@
|
|||||||
|
<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<!--
|
||||||
|
Strategy-as-Code — attīstības plānošanas dokuments kā dati, shēma v0.1
|
||||||
|
PPP Asociācija (PPPA), Valsts PirmKods. Melnraksts.
|
||||||
|
-->
|
||||||
|
<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"
|
||||||
|
xmlns="urn:pppa:vpk:strategija:0.1"
|
||||||
|
xmlns:s="urn:pppa:vpk:strategija:0.1"
|
||||||
|
targetNamespace="urn:pppa:vpk:strategija:0.1"
|
||||||
|
elementFormDefault="qualified" attributeFormDefault="unqualified" version="0.1" xml:lang="lv">
|
||||||
|
|
||||||
|
<xs:annotation>
|
||||||
|
<xs:documentation xml:lang="lv">
|
||||||
|
Attīstības plānošanas dokuments (Attīstības plānošanas sistēmas likums) kā strukturēti dati: nodaļas, numurētie punkti
|
||||||
|
ar to veidu (teksts, mērķis, indikators, uzdevums, telpiskās attīstības virziens), indikatoru vērtības, uzdevumu atbildīgās
|
||||||
|
institūcijas, finansējuma avoti, pamatojuma avoti un zemsvītras piezīmes.
|
||||||
|
|
||||||
|
Teksts un vērtības — kā drukāti avota dokumentā. Institūcijas norāda ar VPK ID no Valsts institūciju reģistra
|
||||||
|
(Valdibas-Deklaracija-as-Code data/organizacijas.xml); drukātais apzīmējums saglabājas atribūtā label.
|
||||||
|
|
||||||
|
Identifikatori:
|
||||||
|
dokuments lv-nap2027
|
||||||
|
nodaļa lv-nap2027.pr1 (prioritāte), lv-nap2027.pr1.rv1 (rīcības virziens; numurs — visā dokumentā),
|
||||||
|
lv-nap2027.pr2.rv6.t1 (tematiska apakšnodaļa), lv-nap2027.introduction u. c.
|
||||||
|
punkts lv-nap2027.i068 — dokumentā drukātais punkta numurs [68]
|
||||||
|
pamatojums lv-nap2027.a001 — pielikuma pamatojuma punkta numurs; nenumurēta rindkopa — lv-nap2027.annex.u1
|
||||||
|
</xs:documentation>
|
||||||
|
</xs:annotation>
|
||||||
|
|
||||||
|
<!-- ============================================================ simple types -->
|
||||||
|
<xs:simpleType name="NonEmptyText"><xs:restriction base="xs:string"><xs:minLength value="1"/></xs:restriction></xs:simpleType>
|
||||||
|
<xs:simpleType name="DocId"><xs:restriction base="xs:string"><xs:pattern value="lv-[a-z0-9]+(-[a-z0-9]+)*"/></xs:restriction></xs:simpleType>
|
||||||
|
<xs:simpleType name="ElementId"><xs:restriction base="xs:string"><xs:pattern value="lv-[a-z0-9]+(-[a-z0-9]+)*(\.[a-z][a-zA-Z]*[0-9]*)+"/></xs:restriction></xs:simpleType>
|
||||||
|
<xs:simpleType name="OrgId">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">VPK ID — SS-IIII no Valsts institūciju reģistra.</xs:documentation></xs:annotation>
|
||||||
|
<xs:restriction base="xs:string"><xs:pattern value="[0-9]{2}-[0-9]{4}"/></xs:restriction>
|
||||||
|
</xs:simpleType>
|
||||||
|
<xs:simpleType name="Sha256"><xs:restriction base="xs:string"><xs:pattern value="[0-9a-f]{64}"/></xs:restriction></xs:simpleType>
|
||||||
|
<xs:simpleType name="FundingCode"><xs:restriction base="xs:string"><xs:pattern value="[A-Z0-9][A-Z0-9-]*"/></xs:restriction></xs:simpleType>
|
||||||
|
|
||||||
|
<xs:simpleType name="DocumentType">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Attīstības plānošanas dokumenta veids (Attīstības plānošanas sistēmas likums, MK noteikumi Nr. 737).</xs:documentation></xs:annotation>
|
||||||
|
<xs:restriction base="xs:string">
|
||||||
|
<xs:enumeration value="ilgtspējīgas attīstības stratēģija"/>
|
||||||
|
<xs:enumeration value="nacionālais attīstības plāns"/>
|
||||||
|
<xs:enumeration value="nacionālās drošības koncepcija"/>
|
||||||
|
<xs:enumeration value="pamatnostādnes"/>
|
||||||
|
<xs:enumeration value="plāns"/>
|
||||||
|
<xs:enumeration value="konceptuāls ziņojums"/>
|
||||||
|
<xs:enumeration value="cits"/>
|
||||||
|
</xs:restriction>
|
||||||
|
</xs:simpleType>
|
||||||
|
|
||||||
|
<xs:simpleType name="SectionKind">
|
||||||
|
<xs:restriction base="xs:string">
|
||||||
|
<xs:enumeration value="introduction"/><xs:enumeration value="vision"/><xs:enumeration value="framework"/>
|
||||||
|
<xs:enumeration value="strategicGoals"/><xs:enumeration value="priority"/><xs:enumeration value="actionLine"/>
|
||||||
|
<xs:enumeration value="theme"/><xs:enumeration value="spatial"/><xs:enumeration value="implementation"/>
|
||||||
|
</xs:restriction>
|
||||||
|
</xs:simpleType>
|
||||||
|
|
||||||
|
<xs:simpleType name="ItemKind">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">
|
||||||
|
text = skaidrojošs teksts; strategicGoal = stratēģiskais mērķis; priorityGoal = prioritātes mērķis;
|
||||||
|
actionLineGoal = rīcības virziena mērķis; indicator = mērķa indikators (tabulas rinda); task = rīcības virziena uzdevums
|
||||||
|
(tabulas rinda); spatialDirection = telpiskās attīstības virziens nacionālo interešu telpā.
|
||||||
|
</xs:documentation></xs:annotation>
|
||||||
|
<xs:restriction base="xs:string">
|
||||||
|
<xs:enumeration value="text"/><xs:enumeration value="strategicGoal"/><xs:enumeration value="priorityGoal"/>
|
||||||
|
<xs:enumeration value="actionLineGoal"/><xs:enumeration value="indicator"/><xs:enumeration value="task"/>
|
||||||
|
<xs:enumeration value="spatialDirection"/>
|
||||||
|
</xs:restriction>
|
||||||
|
</xs:simpleType>
|
||||||
|
|
||||||
|
<xs:simpleType name="ActorMatch">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">
|
||||||
|
direct = tā pati institūcija; renamed = pārdēvēta (reģistrā FormerName); historical = likvidēta vai apvienota (reģistrā validTo un successor);
|
||||||
|
group = institūciju grupa; assumed = drukātais apzīmējums nav saīsinājumu sarakstā, nozīme pieņemta; none = nav reģistrā.
|
||||||
|
</xs:documentation></xs:annotation>
|
||||||
|
<xs:restriction base="xs:string">
|
||||||
|
<xs:enumeration value="direct"/><xs:enumeration value="renamed"/><xs:enumeration value="historical"/>
|
||||||
|
<xs:enumeration value="group"/><xs:enumeration value="assumed"/><xs:enumeration value="none"/>
|
||||||
|
</xs:restriction>
|
||||||
|
</xs:simpleType>
|
||||||
|
|
||||||
|
<!-- ============================================================ root -->
|
||||||
|
<xs:element name="PlanningDocument">
|
||||||
|
<xs:complexType>
|
||||||
|
<xs:sequence>
|
||||||
|
<xs:element name="Metadata" type="MetadataType"/>
|
||||||
|
<xs:element name="Abbreviations" minOccurs="0">
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Abbreviation" maxOccurs="unbounded">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Dokumenta saīsinājumu saraksta ieraksts, kā drukāts; institūcijai — VPK ID.</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:simpleContent><xs:extension base="NonEmptyText">
|
||||||
|
<xs:attribute name="term" type="NonEmptyText" use="required"/>
|
||||||
|
<xs:attribute name="org" type="OrgId"/>
|
||||||
|
</xs:extension></xs:simpleContent></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="Actors">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Visi atbildīgo un līdzatbildīgo institūciju apzīmējumi, kas lietoti dokumentā, un to sasaiste ar VPK ID.</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Actor" maxOccurs="unbounded">
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Note" type="NonEmptyText" minOccurs="0"/>
|
||||||
|
</xs:sequence>
|
||||||
|
<xs:attribute name="label" type="NonEmptyText" use="required"/>
|
||||||
|
<xs:attribute name="org" type="OrgId"/>
|
||||||
|
<xs:attribute name="match" type="ActorMatch" use="required"/>
|
||||||
|
</xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="FundingSources">
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="FundingSource" maxOccurs="unbounded">
|
||||||
|
<xs:complexType><xs:simpleContent><xs:extension base="NonEmptyText">
|
||||||
|
<xs:attribute name="code" type="FundingCode" use="required"/>
|
||||||
|
</xs:extension></xs:simpleContent></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="Section" type="SectionType" maxOccurs="unbounded"/>
|
||||||
|
<xs:element name="Annex" type="AnnexType" minOccurs="0" maxOccurs="unbounded"/>
|
||||||
|
<xs:element name="Footnotes" minOccurs="0">
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Footnote" maxOccurs="unbounded">
|
||||||
|
<xs:complexType><xs:simpleContent><xs:extension base="NonEmptyText">
|
||||||
|
<xs:attribute name="n" type="xs:positiveInteger" use="required"/>
|
||||||
|
<xs:attribute name="page" type="xs:positiveInteger" use="required"/>
|
||||||
|
</xs:extension></xs:simpleContent></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence>
|
||||||
|
<xs:attribute name="id" type="DocId" use="required"/>
|
||||||
|
<xs:attribute name="schemaVersion" type="xs:string" use="required" fixed="0.1"/>
|
||||||
|
</xs:complexType>
|
||||||
|
|
||||||
|
<xs:key name="elementKey"><xs:selector xpath=".//s:Section|.//s:Item|.//s:Evidence"/><xs:field xpath="@id"/></xs:key>
|
||||||
|
<xs:keyref name="indicatorRef" refer="elementKey"><xs:selector xpath=".//s:IndicatorRef"/><xs:field xpath="@ref"/></xs:keyref>
|
||||||
|
<xs:keyref name="evidenceSection" refer="elementKey"><xs:selector xpath=".//s:Evidence"/><xs:field xpath="@section"/></xs:keyref>
|
||||||
|
<xs:key name="itemNumber"><xs:selector xpath=".//s:Item"/><xs:field xpath="@n"/></xs:key>
|
||||||
|
<xs:key name="actorKey"><xs:selector xpath="s:Actors/s:Actor"/><xs:field xpath="@label"/></xs:key>
|
||||||
|
<xs:keyref name="actorRef" refer="actorKey"><xs:selector xpath=".//s:Responsible/s:Actor|.//s:CoResponsible/s:Actor"/><xs:field xpath="@label"/></xs:keyref>
|
||||||
|
<xs:key name="fundingKey"><xs:selector xpath="s:FundingSources/s:FundingSource"/><xs:field xpath="@code"/></xs:key>
|
||||||
|
<xs:keyref name="fundingRef" refer="fundingKey"><xs:selector xpath=".//s:Funding/s:Source"/><xs:field xpath="@code"/></xs:keyref>
|
||||||
|
<xs:key name="footnoteKey"><xs:selector xpath="s:Footnotes/s:Footnote"/><xs:field xpath="@n"/></xs:key>
|
||||||
|
<xs:keyref name="footnoteRef" refer="footnoteKey"><xs:selector xpath=".//s:FootnoteRef"/><xs:field xpath="@n"/></xs:keyref>
|
||||||
|
</xs:element>
|
||||||
|
|
||||||
|
<!-- ============================================================ metadata -->
|
||||||
|
<xs:complexType name="MetadataType">
|
||||||
|
<xs:sequence>
|
||||||
|
<xs:element name="Title" type="NonEmptyText"/>
|
||||||
|
<xs:element name="ShortTitle" type="NonEmptyText" minOccurs="0"/>
|
||||||
|
<xs:element name="DocumentType" type="DocumentType"/>
|
||||||
|
<xs:element name="Period">
|
||||||
|
<xs:complexType><xs:attribute name="from" type="xs:gYear" use="required"/><xs:attribute name="to" type="xs:gYear" use="required"/></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="Approval">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Lēmums, ar kuru dokuments apstiprināts.</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Body" type="NonEmptyText"/>
|
||||||
|
<xs:element name="Act" type="NonEmptyText"/>
|
||||||
|
<xs:element name="Number" type="NonEmptyText" minOccurs="0"/>
|
||||||
|
<xs:element name="Date" type="xs:date"/>
|
||||||
|
<xs:element name="Url" type="xs:anyURI" minOccurs="0"/>
|
||||||
|
</xs:sequence></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="Developer" minOccurs="0">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Dokumenta izstrādātājs, kā norādīts titullapā.</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:simpleContent><xs:extension base="NonEmptyText"><xs:attribute name="org" type="OrgId"/></xs:extension></xs:simpleContent></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="Source">
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Url" type="xs:anyURI"/>
|
||||||
|
<xs:element name="File" type="NonEmptyText"/>
|
||||||
|
<xs:element name="SHA256" type="Sha256"/>
|
||||||
|
<xs:element name="Pages" type="xs:positiveInteger"/>
|
||||||
|
<xs:element name="Retrieved" type="xs:date"/>
|
||||||
|
</xs:sequence></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="Conversion">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Kas un kā dokumentu pārveidoja datos. Saistošs ir apstiprinātais avota dokuments.</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="By" type="NonEmptyText"/>
|
||||||
|
<xs:element name="Method" type="NonEmptyText"/>
|
||||||
|
</xs:sequence></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="DataVersion" maxOccurs="unbounded">
|
||||||
|
<xs:complexType><xs:sequence><xs:element name="Change" type="NonEmptyText" maxOccurs="unbounded"/></xs:sequence>
|
||||||
|
<xs:attribute name="number" type="xs:positiveInteger" use="required"/>
|
||||||
|
<xs:attribute name="date" type="xs:date" use="required"/>
|
||||||
|
</xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence>
|
||||||
|
</xs:complexType>
|
||||||
|
|
||||||
|
<!-- ============================================================ structure -->
|
||||||
|
<xs:complexType name="SectionType">
|
||||||
|
<xs:sequence>
|
||||||
|
<xs:element name="Title" type="NonEmptyText"/>
|
||||||
|
<xs:element name="IndicativeFunding" minOccurs="0">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Rīcības virziena pasākumu īstenošanai indikatīvi pieejamais finanšu apjoms: teksts kā drukāts; millionEur — skaitlis (milj. EUR).</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:simpleContent><xs:extension base="NonEmptyText">
|
||||||
|
<xs:attribute name="millionEur" type="xs:decimal" use="required"/>
|
||||||
|
</xs:extension></xs:simpleContent></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="Note" minOccurs="0" maxOccurs="unbounded">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Tabulas piezīme („*”) vai lodziņš; area — telpiskās perspektīvas tabulas rinda, pie kuras tas drukāts.</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:simpleContent><xs:extension base="NonEmptyText">
|
||||||
|
<xs:attribute name="area" type="NonEmptyText"/>
|
||||||
|
</xs:extension></xs:simpleContent></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:choice minOccurs="0" maxOccurs="unbounded">
|
||||||
|
<xs:element name="Item" type="ItemType"/>
|
||||||
|
<xs:element name="Section" type="SectionType"/>
|
||||||
|
</xs:choice>
|
||||||
|
</xs:sequence>
|
||||||
|
<xs:attribute name="id" type="ElementId" use="required"/>
|
||||||
|
<xs:attribute name="kind" type="SectionKind" use="required"/>
|
||||||
|
<xs:attribute name="n" type="xs:positiveInteger"/>
|
||||||
|
</xs:complexType>
|
||||||
|
|
||||||
|
<xs:complexType name="ItemType">
|
||||||
|
<xs:sequence>
|
||||||
|
<xs:element name="Title" type="NonEmptyText" minOccurs="0">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Treknrakstā drukātais punkta sākums (stratēģiskajiem mērķiem — mērķa nosaukums).</xs:documentation></xs:annotation>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="Area" type="NonEmptyText" minOccurs="0">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Nacionālo interešu telpa (telpiskās attīstības virzienam).</xs:documentation></xs:annotation>
|
||||||
|
</xs:element>
|
||||||
|
<xs:choice>
|
||||||
|
<xs:element name="Text" type="NonEmptyText"/>
|
||||||
|
<xs:element name="Indicator" type="IndicatorType"/>
|
||||||
|
<xs:element name="Task" type="TaskType"/>
|
||||||
|
</xs:choice>
|
||||||
|
<xs:element name="FootnoteRef" minOccurs="0" maxOccurs="unbounded">
|
||||||
|
<xs:complexType><xs:attribute name="n" type="xs:positiveInteger" use="required"/></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence>
|
||||||
|
<xs:attribute name="id" type="ElementId" use="required"/>
|
||||||
|
<xs:attribute name="n" type="xs:positiveInteger" use="required"><xs:annotation><xs:documentation xml:lang="lv">Drukātais punkta numurs.</xs:documentation></xs:annotation></xs:attribute>
|
||||||
|
<xs:attribute name="kind" type="ItemKind" use="required"/>
|
||||||
|
<xs:attribute name="page" type="xs:positiveInteger" use="required"/>
|
||||||
|
</xs:complexType>
|
||||||
|
|
||||||
|
<xs:complexType name="IndicatorType">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Mērķa indikatora tabulas rinda. Vērtības — kā drukātas (var būt „~27 %”, „1,5/9,2”, teksts). Saliktam indikatoram (piem., vairākas apakšrindas) vērtības drukātā secībā, atdalītas ar atstarpi.</xs:documentation></xs:annotation>
|
||||||
|
<xs:sequence>
|
||||||
|
<xs:element name="Name" type="NonEmptyText"/>
|
||||||
|
<xs:element name="Unit" type="NonEmptyText" minOccurs="0"/>
|
||||||
|
<xs:element name="BaseYear" type="NonEmptyText" minOccurs="0"/>
|
||||||
|
<xs:element name="BaseValue" type="NonEmptyText" minOccurs="0"/>
|
||||||
|
<xs:element name="Target" minOccurs="0" maxOccurs="unbounded">
|
||||||
|
<xs:complexType><xs:simpleContent><xs:extension base="NonEmptyText">
|
||||||
|
<xs:attribute name="year" type="xs:gYear" use="required"/>
|
||||||
|
</xs:extension></xs:simpleContent></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="DataSource" type="NonEmptyText" minOccurs="0"/>
|
||||||
|
</xs:sequence>
|
||||||
|
</xs:complexType>
|
||||||
|
|
||||||
|
<xs:complexType name="TaskType">
|
||||||
|
<xs:sequence>
|
||||||
|
<xs:element name="Text" type="NonEmptyText"/>
|
||||||
|
<xs:element name="Responsible" type="ActorListType"/>
|
||||||
|
<xs:element name="CoResponsible" type="ActorListType" minOccurs="0"/>
|
||||||
|
<xs:element name="Funding" minOccurs="0">
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Source" maxOccurs="unbounded">
|
||||||
|
<xs:complexType><xs:attribute name="code" type="FundingCode" use="required"/><xs:attribute name="label" type="NonEmptyText" use="required"/></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence>
|
||||||
|
<xs:attribute name="printed" type="NonEmptyText" use="required"/>
|
||||||
|
</xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
<xs:element name="TaskIndicator" minOccurs="0" maxOccurs="unbounded">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Uzdevuma tabulas aile „Indikators”: ja nosaukums atbilst dokumenta mērķa indikatoram — IndicatorRef; citādi uzdevuma līmeņa indikators tikai ar nosaukumu.</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Name" type="NonEmptyText"/>
|
||||||
|
<xs:element name="IndicatorRef" minOccurs="0">
|
||||||
|
<xs:complexType>
|
||||||
|
<xs:attribute name="ref" type="ElementId" use="required"/>
|
||||||
|
<xs:attribute name="match" use="required">
|
||||||
|
<xs:simpleType><xs:restriction base="xs:string">
|
||||||
|
<xs:enumeration value="exact"/><xs:enumeration value="shortened"/><xs:enumeration value="similar"/>
|
||||||
|
</xs:restriction></xs:simpleType>
|
||||||
|
</xs:attribute>
|
||||||
|
</xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence>
|
||||||
|
</xs:complexType>
|
||||||
|
|
||||||
|
<xs:complexType name="ActorListType">
|
||||||
|
<xs:sequence>
|
||||||
|
<xs:element name="Actor" maxOccurs="unbounded">
|
||||||
|
<xs:complexType>
|
||||||
|
<xs:attribute name="label" type="NonEmptyText" use="required"/>
|
||||||
|
<xs:attribute name="org" type="OrgId"/>
|
||||||
|
<xs:attribute name="qualifier" type="NonEmptyText"/>
|
||||||
|
</xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence>
|
||||||
|
<xs:attribute name="printed" type="NonEmptyText" use="required"><xs:annotation><xs:documentation xml:lang="lv">Ailes teksts, kā drukāts.</xs:documentation></xs:annotation></xs:attribute>
|
||||||
|
</xs:complexType>
|
||||||
|
|
||||||
|
<xs:complexType name="AnnexType">
|
||||||
|
<xs:sequence>
|
||||||
|
<xs:element name="Title" type="NonEmptyText"/>
|
||||||
|
<xs:element name="Evidence" maxOccurs="unbounded">
|
||||||
|
<xs:annotation><xs:documentation xml:lang="lv">Pamatojuma punkts: teksts kā drukāts, tajā minētās tīmekļa adreses; section — prioritāte vai rīcības virziens, kuru pamato;
|
||||||
|
Problem — treknrakstā drukātā problēma, pie kuras punkts grupēts. Nenumurētai rindkopai n nav, identifikators — {dokuments}.annex.uN.</xs:documentation></xs:annotation>
|
||||||
|
<xs:complexType><xs:sequence>
|
||||||
|
<xs:element name="Problem" type="NonEmptyText" minOccurs="0"/>
|
||||||
|
<xs:element name="Text" type="NonEmptyText"/>
|
||||||
|
<xs:element name="Url" type="xs:anyURI" minOccurs="0" maxOccurs="unbounded"/>
|
||||||
|
<xs:element name="FootnoteRef" minOccurs="0" maxOccurs="unbounded">
|
||||||
|
<xs:complexType><xs:attribute name="n" type="xs:positiveInteger" use="required"/></xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence>
|
||||||
|
<xs:attribute name="id" type="ElementId" use="required"/>
|
||||||
|
<xs:attribute name="n" type="xs:positiveInteger"/>
|
||||||
|
<xs:attribute name="section" type="ElementId" use="required"/>
|
||||||
|
<xs:attribute name="page" type="xs:positiveInteger" use="required"/>
|
||||||
|
</xs:complexType>
|
||||||
|
</xs:element>
|
||||||
|
</xs:sequence>
|
||||||
|
<xs:attribute name="id" type="ElementId" use="required"/>
|
||||||
|
</xs:complexType>
|
||||||
|
</xs:schema>
|
||||||
101
sources/nap2027/dalibnieki.yaml
Normal file
101
sources/nap2027/dalibnieki.yaml
Normal file
@@ -0,0 +1,101 @@
|
|||||||
|
# NAP2027 uzdevumu tabulās drukātie atbildīgo un līdzatbildīgo institūciju apzīmējumi → VPK ID
|
||||||
|
# (Valdibas-Deklaracija-as-Code data/organizacijas.xml).
|
||||||
|
#
|
||||||
|
# match:
|
||||||
|
# direct — tā pati institūcija ar to pašu nosaukumu
|
||||||
|
# renamed — tā pati institūcija, kopš NAP2027 pieņemšanas pārdēvēta (reģistrā FormerName)
|
||||||
|
# historical — institūcija likvidēta vai apvienota; ID ar validTo un successor
|
||||||
|
# group — institūciju grupa (resors 00) vai pašvaldības
|
||||||
|
# assumed — drukāts apzīmējums, kas nav saīsinājumu sarakstā; nozīme pieņemta, jāapstiprina
|
||||||
|
# none — nav valsts institūcija vai grupa reģistrā; ID nepiešķir
|
||||||
|
#
|
||||||
|
# Dokumentā saglabā drukāto apzīmējumu; ID ir sasaiste, nevis aizstājējs.
|
||||||
|
actors:
|
||||||
|
- {label: "AizM", org: "10-0000", match: direct}
|
||||||
|
- {label: "ĀM", org: "11-0000", match: direct}
|
||||||
|
- {label: "EM", org: "12-0000", match: direct}
|
||||||
|
- {label: "FM", org: "13-0000", match: direct}
|
||||||
|
- {label: "IeM", org: "14-0000", match: direct}
|
||||||
|
- {label: "IZM", org: "15-0000", match: direct}
|
||||||
|
- {label: "ZM", org: "16-0000", match: direct}
|
||||||
|
- {label: "SM", org: "17-0000", match: direct}
|
||||||
|
- {label: "LM", org: "18-0000", match: direct}
|
||||||
|
- {label: "TM", org: "19-0000", match: direct}
|
||||||
|
- {label: "VARAM", org: "21-0000", match: renamed,
|
||||||
|
note: "NAP2027 pieņemšanas brīdī — Vides aizsardzības un reģionālās attīstības ministrija; no 2024-07-01 Viedās administrācijas un reģionālās attīstības ministrija. Vides aizsardzības politika līdz 2024-06-30 nodota Klimata un enerģētikas ministrijai (20-0000), dabas aizsardzība palika VARAM."}
|
||||||
|
- {label: "KM", org: "22-0000", match: direct}
|
||||||
|
- {label: "VM", org: "29-0000", match: direct}
|
||||||
|
- {label: "VK", org: "03-0000", match: direct}
|
||||||
|
- {label: "KNAB", org: "04-0000", match: direct}
|
||||||
|
- {label: "SIF", org: "08-0000", match: direct}
|
||||||
|
- {label: "SPRK", org: "09-0000", match: direct}
|
||||||
|
- {label: "FID", org: "14-0693", match: direct}
|
||||||
|
- {label: "NEPLP", org: "47-0000", match: direct}
|
||||||
|
- {label: "Prokuratūra", org: "32-0000", match: direct}
|
||||||
|
- {label: "PKC", org: "03-9001", match: historical,
|
||||||
|
note: "Pārresoru koordinācijas centrs likvidēts; funkcijas no 2023-03-01 pārņēma Valsts kanceleja (03-0000)."}
|
||||||
|
- {label: "PKC (DLC)", org: "03-9001", match: historical,
|
||||||
|
note: "DLC — Demogrāfisko lietu centrs, PKC struktūrvienība; funkcijas no 2023-03-01 Valsts kancelejā (03-0000)."}
|
||||||
|
- {label: "FKTK", org: "94-0002", match: historical,
|
||||||
|
note: "Finanšu un kapitāla tirgus komisija no 2023-01-01 pievienota Latvijas Bankai (94-0001)."}
|
||||||
|
- {label: "Visas ministrijas", org: "00-0002", match: group}
|
||||||
|
- {label: "visas ministrijas", org: "00-0002", match: group}
|
||||||
|
- {label: "visas citas ministrijas", org: "00-0002", match: group, note: "Visas ministrijas, izņemot atbildīgo."}
|
||||||
|
- {label: "resori", org: "00-0002", match: group, note: "Drukāts „resori, kuri sniedz pakalpojumus VPVKAC” — ministrijas ar padotības iestādēm, kuru pakalpojumi pieejami VPVKAC."}
|
||||||
|
- {label: "Pašvaldības", org: "00-0003", match: group}
|
||||||
|
- {label: "pašvaldības", org: "00-0003", match: group}
|
||||||
|
- {label: "NVO", org: "00-0005", match: group}
|
||||||
|
- {label: "Tiesas", org: "00-0006", match: group}
|
||||||
|
- {label: "plānošanas reģioni", org: "00-0007", match: group}
|
||||||
|
- {label: "VKS", org: "00-0008", match: group}
|
||||||
|
- {label: "augstskolas", org: "00-0009", match: group}
|
||||||
|
- {label: "augstākās izglītības iestādes", org: "00-0009", match: group}
|
||||||
|
- {label: "izglītības iestādes", org: "00-0010", match: group}
|
||||||
|
- {label: "reliģiskās organizācijas", org: "00-0011", match: group}
|
||||||
|
- {label: "komersanti", org: "00-0012", match: group}
|
||||||
|
- {label: "LTRK", org: "93-0001", match: direct}
|
||||||
|
- {label: "LTKR", org: "93-0001", match: assumed, note: "Drukāts „LTKR”; saīsinājumu sarakstā ir LTRK — pieņemta drukas kļūda."}
|
||||||
|
- {label: "LDDK", org: "93-0002", match: direct}
|
||||||
|
- {label: "LPS", org: "93-0003", match: direct}
|
||||||
|
- {label: "LBAS", org: "93-0004", match: direct}
|
||||||
|
- {label: "LABS", org: "93-0004", match: assumed, note: "Drukāts „LABS”; saīsinājumu sarakstā tāda nav. Pieņemts LBAS (tajā pašā uzdevumā arī LDDK) — jāapstiprina."}
|
||||||
|
- {label: "LIZDA", org: "93-0005", match: assumed, note: "Nav saīsinājumu sarakstā; Latvijas Izglītības un zinātnes darbinieku arodbiedrība."}
|
||||||
|
- {label: "LLPA", org: "93-0006", match: assumed, note: "Nav saīsinājumu sarakstā; Latvijas Lielo pilsētu asociācija."}
|
||||||
|
- {label: "Altum", org: "91-0004", match: direct}
|
||||||
|
- {label: "VSIA “Latvijas Vēstnesis”", org: "91-0005", match: direct}
|
||||||
|
- {label: "Nasdaq", org: null, match: none, note: "Nasdaq Riga — privāta biržas sabiedrība, nav valsts institūcija."}
|
||||||
|
# Apzīmējumi, ko tabulā drukā bez komata starp divām institūcijām
|
||||||
|
split:
|
||||||
|
"LBAS plānošanas reģioni": ["LBAS", "plānošanas reģioni"]
|
||||||
|
"pašvaldības NVO": ["pašvaldības", "NVO"]
|
||||||
|
# Teksta daļa, kas pēc komata turpina iepriekšējo apzīmējumu
|
||||||
|
qualifiers:
|
||||||
|
"kuri sniedz pakalpojumus VPVKAC": "resori"
|
||||||
|
|
||||||
|
# Saīsinājumu saraksta (3.–4. lpp.) institūcijas, kas uzdevumu tabulās nav atbildīgās
|
||||||
|
abbreviations:
|
||||||
|
CSP: "12-0039"
|
||||||
|
DAP: "21-0650"
|
||||||
|
LAD: "16-0323"
|
||||||
|
LB: "94-0001"
|
||||||
|
LIAA: "12-0045"
|
||||||
|
LVĢMC: "91-0006"
|
||||||
|
NVD: "29-0674"
|
||||||
|
RD: "90-0004"
|
||||||
|
SPKC: "29-0677"
|
||||||
|
VID: "13-0056"
|
||||||
|
VSAA: "18-0453"
|
||||||
|
|
||||||
|
# Finansējuma avoti, kā drukāti uzdevumu tabulās
|
||||||
|
funding:
|
||||||
|
"VB": VB
|
||||||
|
"ES fondi": ES-FONDI
|
||||||
|
"citi finanšu avoti": CITI
|
||||||
|
"pašvaldības": PASV
|
||||||
|
"Horizon Europe": HORIZON
|
||||||
|
"Digital Europe": DIGITAL
|
||||||
|
"Urban Europe": URBAN
|
||||||
|
"EEZ/ Norvēģijas finanšu instruments": EEZ-NFI
|
||||||
|
"Šveices programma": CH
|
||||||
|
"ES programmas jaunatnes jomā": ES-JAUNATNE
|
||||||
|
"pilsoniskā iniciatīva": PILSONISKA
|
||||||
257
tools/build_nap2027.py
Normal file
257
tools/build_nap2027.py
Normal file
@@ -0,0 +1,257 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""NAP2027: starpposma JSON (parse_nap.py) + institūciju sasaiste (sources/nap2027/dalibnieki.yaml) → data/nap2027/lv-nap2027.xml
|
||||||
|
|
||||||
|
python3 tools/parse_nap.py sources/nap2027/NAP2027.pdf build/nap2027.json
|
||||||
|
python3 tools/build_nap2027.py build/nap2027.json data/nap2027/lv-nap2027.xml
|
||||||
|
"""
|
||||||
|
import datetime as dt
|
||||||
|
import difflib
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
from xml.sax.saxutils import escape, quoteattr
|
||||||
|
|
||||||
|
import yaml
|
||||||
|
|
||||||
|
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
DOC = "lv-nap2027"
|
||||||
|
PDF = "sources/nap2027/NAP2027.pdf"
|
||||||
|
SOURCE_URL = "https://www.mk.gov.lv/lv/media/15162/download"
|
||||||
|
RETRIEVED = "2026-10-11"
|
||||||
|
FUNDING_NAMES = {
|
||||||
|
"VB": "Valsts budžets", "ES-FONDI": "Eiropas Savienības fondi", "CITI": "Citi finanšu avoti", "PASV": "Pašvaldību budžeti",
|
||||||
|
"HORIZON": "Horizon Europe", "DIGITAL": "Digital Europe", "URBAN": "Urban Europe",
|
||||||
|
"EEZ-NFI": "EEZ un Norvēģijas finanšu instruments", "CH": "Šveices programma",
|
||||||
|
"ES-JAUNATNE": "ES programmas jaunatnes jomā", "PILSONISKA": "Pilsoniskā iniciatīva",
|
||||||
|
}
|
||||||
|
SUBS = str.maketrans("₀₁₂₃₄₅₆₇₈₉", "0123456789")
|
||||||
|
|
||||||
|
|
||||||
|
def sq(s):
|
||||||
|
return re.sub(r"[^0-9a-zāčēģīķļņšūž%]", "", (s or "").lower().translate(SUBS))
|
||||||
|
|
||||||
|
|
||||||
|
def a(name, value):
|
||||||
|
return f" {name}={quoteattr(str(value))}" if value not in (None, "") else ""
|
||||||
|
|
||||||
|
|
||||||
|
def el(tag, text, **attrs):
|
||||||
|
if text in (None, ""):
|
||||||
|
return ""
|
||||||
|
return f"<{tag}{''.join(a(k, v) for k, v in attrs.items())}>{escape(str(text))}</{tag}>"
|
||||||
|
|
||||||
|
|
||||||
|
def main(src, out):
|
||||||
|
d = json.load(open(src, encoding="utf-8"))
|
||||||
|
m = yaml.safe_load(open(os.path.join(ROOT, "sources/nap2027/dalibnieki.yaml"), encoding="utf-8"))
|
||||||
|
actors = {x["label"]: x for x in m["actors"]}
|
||||||
|
used_actors, used_funding, unresolved = {}, {}, []
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- actors
|
||||||
|
def actor_list(printed):
|
||||||
|
if not printed:
|
||||||
|
return []
|
||||||
|
parts = [p.strip() for p in re.split(r",\s*", printed) if p.strip()]
|
||||||
|
out_ = []
|
||||||
|
for p in parts:
|
||||||
|
if p in m.get("qualifiers", {}):
|
||||||
|
if out_ and out_[-1]["label"] == m["qualifiers"][p]:
|
||||||
|
out_[-1]["qualifier"] = p
|
||||||
|
continue
|
||||||
|
for q in m.get("split", {}).get(p, [p]):
|
||||||
|
if q not in actors:
|
||||||
|
unresolved.append(q)
|
||||||
|
continue
|
||||||
|
used_actors[q] = actors[q]
|
||||||
|
out_.append({"label": q, "org": actors[q].get("org")})
|
||||||
|
return out_
|
||||||
|
|
||||||
|
def actors_xml(tag, printed):
|
||||||
|
lst = actor_list(printed)
|
||||||
|
if not lst:
|
||||||
|
return ""
|
||||||
|
inner = "".join(f"<Actor{a('label', x['label'])}{a('org', x.get('org'))}{a('qualifier', x.get('qualifier'))}/>" for x in lst)
|
||||||
|
return f"<{tag}{a('printed', printed)}>{inner}</{tag}>"
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- funding
|
||||||
|
fmap = m["funding"]
|
||||||
|
|
||||||
|
def funding_xml(printed):
|
||||||
|
if not printed:
|
||||||
|
return ""
|
||||||
|
srcs = []
|
||||||
|
for p in [p.strip() for p in re.split(r",\s*", printed) if p.strip()]:
|
||||||
|
labels = [p] if p in fmap else [x.strip() for x in p.split("/")]
|
||||||
|
for lab in labels:
|
||||||
|
if lab not in fmap:
|
||||||
|
unresolved.append("finansējums: " + lab)
|
||||||
|
continue
|
||||||
|
used_funding[fmap[lab]] = True
|
||||||
|
srcs.append(f"<Source{a('code', fmap[lab])}{a('label', lab)}/>")
|
||||||
|
return f"<Funding{a('printed', printed)}>{''.join(srcs)}</Funding>" if srcs else ""
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- task indicator → document indicator
|
||||||
|
inds = [x for x in d["items"] if x["kind"] == "indicator"]
|
||||||
|
rv_of = lambda sec: ".".join(sec.split(".")[:2])
|
||||||
|
match_stats = {"exact": 0, "shortened": 0, "similar": 0, "split": 0, "task-level": 0}
|
||||||
|
|
||||||
|
def best_indicator(name, sec):
|
||||||
|
s = sq(name)
|
||||||
|
if len(s) < 4:
|
||||||
|
return None, None
|
||||||
|
same = [i for i in inds if rv_of(i["section"]) == rv_of(sec)]
|
||||||
|
for pool in (same, inds):
|
||||||
|
for i in pool:
|
||||||
|
if sq(i["name"]) == s:
|
||||||
|
return i, "exact"
|
||||||
|
for i in pool:
|
||||||
|
t = sq(i["name"])
|
||||||
|
head = sq(re.split(r"[(,]", i["name"])[0]) # name without the bracketed or comma explanation
|
||||||
|
if len(s) >= 10 and t.startswith(s) and (len(s) >= 0.6 * len(t) or head == s):
|
||||||
|
return i, "shortened" # printed without the bracketed explanation, or slightly shortened
|
||||||
|
r, i = max(((difflib.SequenceMatcher(None, s, sq(i["name"])).ratio() + (0.02 if i in same else 0), i) for i in inds),
|
||||||
|
key=lambda z: z[0])
|
||||||
|
return (i, "similar") if r >= 0.85 else (None, None)
|
||||||
|
|
||||||
|
def split_merged(name, sec):
|
||||||
|
"""Several indicator names printed without an empty line between them: split before capitalised words
|
||||||
|
(not all-caps abbreviations); keep matched pieces separate, join unmatched neighbours back together."""
|
||||||
|
w = name.split(" ")
|
||||||
|
cuts = [k for k in range(1, len(w)) if w[k][:1].isupper() and not w[k].isupper() and not w[k - 1].endswith(("(", "–", "-", "/"))]
|
||||||
|
if not cuts:
|
||||||
|
return None
|
||||||
|
segs = [" ".join(w[i:j]) for i, j in zip([0] + cuts, cuts + [len(w)])]
|
||||||
|
res = []
|
||||||
|
for sgm in segs:
|
||||||
|
i, how = best_indicator(sgm, sec)
|
||||||
|
if i and how in ("exact", "shortened"):
|
||||||
|
res.append((sgm, i, how))
|
||||||
|
elif res and res[-1][1] is None:
|
||||||
|
res[-1] = (res[-1][0] + " " + sgm, None, None)
|
||||||
|
else:
|
||||||
|
res.append((sgm, None, None))
|
||||||
|
return res if any(r[1] for r in res) and len(res) > 1 else None
|
||||||
|
|
||||||
|
def task_indicators(t):
|
||||||
|
xs = []
|
||||||
|
for nm in t.get("indicatorNames") or []:
|
||||||
|
i, how = best_indicator(nm, t["section"])
|
||||||
|
parts = [(nm, i, how)] if i else (split_merged(nm, t["section"]) or [(nm, None, None)])
|
||||||
|
if len(parts) > 1:
|
||||||
|
match_stats["split"] += 1
|
||||||
|
for name, i, how in parts:
|
||||||
|
match_stats[how or "task-level"] += 1
|
||||||
|
ref = f"<IndicatorRef{a('ref', DOC + '.i%03d' % i['number'])}{a('match', how)}/>" if i else ""
|
||||||
|
xs.append(f"<TaskIndicator>{el('Name', name)}{ref}</TaskIndicator>")
|
||||||
|
return "".join(xs)
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- items
|
||||||
|
def item_xml(x):
|
||||||
|
iid = f"{DOC}.i{x['number']:03d}"
|
||||||
|
body = ""
|
||||||
|
if x["kind"] == "indicator":
|
||||||
|
body = ("<Indicator>" + el("Name", x.get("name")) + el("Unit", x.get("unit")) + el("BaseYear", x.get("baseYear"))
|
||||||
|
+ el("BaseValue", x.get("baseValue")) + el("Target", x.get("target2024"), year="2024")
|
||||||
|
+ el("Target", x.get("target2027"), year="2027") + el("DataSource", x.get("source")) + "</Indicator>")
|
||||||
|
elif x["kind"] == "task":
|
||||||
|
body = ("<Task>" + el("Text", x["text"]) + actors_xml("Responsible", x.get("responsible"))
|
||||||
|
+ actors_xml("CoResponsible", x.get("coResponsible")) + funding_xml(x.get("funding"))
|
||||||
|
+ task_indicators(x) + "</Task>")
|
||||||
|
else:
|
||||||
|
title = x.get("title") if x["kind"] == "strategicGoal" else None
|
||||||
|
body = el("Title", title) + el("Area", x.get("area")) + el("Text", x["text"])
|
||||||
|
refs = "".join(f'<FootnoteRef n="{n}"/>' for n in x.get("footnoteRefs", []))
|
||||||
|
return f'<Item id="{iid}" n="{x["number"]}" kind="{x["kind"]}" page="{x["page"]}">{body}{refs}</Item>'
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- sections (tree)
|
||||||
|
secs = [s for s in d["sections"] if s["kind"] not in ("front", "annex")]
|
||||||
|
by_parent = {}
|
||||||
|
for s in secs:
|
||||||
|
by_parent.setdefault(s.get("parent"), []).append(s)
|
||||||
|
items_by_sec = {}
|
||||||
|
for x in d["items"]:
|
||||||
|
items_by_sec.setdefault(x["section"], []).append(x)
|
||||||
|
|
||||||
|
def sec_num(s):
|
||||||
|
mm = re.search(r"(\d+)$", s["id"])
|
||||||
|
return mm.group(1) if s["kind"] in ("priority", "actionLine", "theme") and mm else None
|
||||||
|
|
||||||
|
def section_xml(s):
|
||||||
|
sid = f"{DOC}.{s['id']}"
|
||||||
|
parts = [el("Title", s["title"])]
|
||||||
|
f = s.get("funding")
|
||||||
|
if f and f.get("millionEur"):
|
||||||
|
parts.append(el("IndicativeFunding", f["text"], millionEur=f["millionEur"].replace(",", ".")))
|
||||||
|
for nt in s.get("notes", []):
|
||||||
|
if isinstance(nt, dict):
|
||||||
|
parts.append(el("Note", nt["text"], area=nt.get("area")))
|
||||||
|
else:
|
||||||
|
parts.append(el("Note", nt))
|
||||||
|
# items and sub-sections in document order (by first item number)
|
||||||
|
children = [(x["number"], item_xml(x)) for x in items_by_sec.get(s["id"], [])]
|
||||||
|
for c in by_parent.get(s["id"], []):
|
||||||
|
first = min([x["number"] for x in d["items"] if x["section"] == c["id"] or x["section"].startswith(c["id"] + ".")] or [10 ** 6])
|
||||||
|
children.append((first - 0.5, section_xml(c)))
|
||||||
|
parts += [c for _, c in sorted(children, key=lambda z: z[0])]
|
||||||
|
return f'<Section id="{sid}" kind="{s["kind"]}"{a("n", sec_num(s))}>' + "".join(parts) + "</Section>"
|
||||||
|
|
||||||
|
top = [s for s in by_parent.get(None, [])]
|
||||||
|
body = "".join(section_xml(s) for s in top)
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- annex
|
||||||
|
un = iter(range(1, 1000))
|
||||||
|
ev = "".join(
|
||||||
|
(f'<Evidence id="{DOC}.a{e["number"]:03d}" n="{e["number"]}"' if e["number"] else f'<Evidence id="{DOC}.annex.u{next(un)}"')
|
||||||
|
+ f' section="{DOC}.{e["section"]}" page="{e["page"]}">'
|
||||||
|
+ el("Problem", (e.get("problem") or "").rstrip(":")) + el("Text", e["text"]) + "".join(el("Url", u) for u in e.get("urls", []))
|
||||||
|
+ "".join(f'<FootnoteRef n="{n}"/>' for n in e.get("footnoteRefs", [])) + "</Evidence>"
|
||||||
|
for e in d["evidence"])
|
||||||
|
annex = f'<Annex id="{DOC}.annex">{el("Title", "NAP2027 prioritāšu pamatojuma avoti")}{ev}</Annex>'
|
||||||
|
notes = "<Footnotes>" + "".join(el("Footnote", n["text"], n=n["number"], page=n["page"]) for n in d["footnotes"]) + "</Footnotes>"
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- head
|
||||||
|
abbr_org = {**{x["label"]: x.get("org") for x in m["actors"] if x["match"] in ("direct", "renamed", "historical")},
|
||||||
|
**m.get("abbreviations", {})}
|
||||||
|
abbrs = "<Abbreviations>" + "".join(el("Abbreviation", x["meaning"], term=x["abbr"], org=abbr_org.get(x["abbr"]))
|
||||||
|
for x in d["abbreviations"]) + "</Abbreviations>"
|
||||||
|
act = "<Actors>" + "".join(
|
||||||
|
f"<Actor{a('label', x['label'])}{a('org', x.get('org'))}{a('match', x['match'])}>" + el("Note", x.get("note")) + "</Actor>"
|
||||||
|
for lab, x in sorted(used_actors.items(), key=lambda z: (z[1].get("org") or "99", z[0]))) + "</Actors>"
|
||||||
|
fund = "<FundingSources>" + "".join(el("FundingSource", FUNDING_NAMES[c], code=c) for c in FUNDING_NAMES if c in used_funding) + "</FundingSources>"
|
||||||
|
sha = hashlib.sha256(open(os.path.join(ROOT, PDF), "rb").read()).hexdigest()
|
||||||
|
meta = ("<Metadata>"
|
||||||
|
+ el("Title", "Latvijas Nacionālais attīstības plāns 2021.–2027. gadam") + el("ShortTitle", "NAP2027")
|
||||||
|
+ el("DocumentType", "nacionālais attīstības plāns") + '<Period from="2021" to="2027"/>'
|
||||||
|
+ "<Approval>" + el("Body", "Latvijas Republikas Saeima") + el("Act", "Saeimas lēmums") + el("Number", "418/Lm13")
|
||||||
|
+ el("Date", "2020-07-02") + "</Approval>"
|
||||||
|
+ el("Developer", "Pārresoru koordinācijas centrs", org="03-9001")
|
||||||
|
+ "<Source>" + el("Url", SOURCE_URL) + el("File", PDF) + el("SHA256", sha) + el("Pages", 127) + el("Retrieved", RETRIEVED) + "</Source>"
|
||||||
|
+ "<Conversion>" + el("By", "PPP Asociācija (PPPA), Valsts PirmKods") + el("Method",
|
||||||
|
"Automātiska nolasīšana no PDF (tools/parse_nap.py: vārdi ar koordinātām, tabulu ailes pēc atstarpēm starp ailēm) un "
|
||||||
|
"institūciju sasaiste pēc sources/nap2027/dalibnieki.yaml (tools/build_nap2027.py); pārbaude — tools/validate.py.") + "</Conversion>"
|
||||||
|
+ f'<DataVersion number="1" date="{dt.date.today().isoformat()}">'
|
||||||
|
+ el("Change", f"Pirmā versija: {len(d['items'])} numurētie punkti [1]–[475], {len(d['evidence'])} pamatojuma punkti, "
|
||||||
|
f"{len(d['footnotes'])} zemsvītras piezīmes, {len(d['abbreviations'])} saīsinājumi.")
|
||||||
|
+ el("Change", "Atbildīgās un līdzatbildīgās institūcijas sasaistītas ar VPK ID; drukātais apzīmējums saglabāts.")
|
||||||
|
+ "</DataVersion></Metadata>")
|
||||||
|
|
||||||
|
xml = ('<?xml version="1.0" encoding="UTF-8"?>\n'
|
||||||
|
f'<PlanningDocument xmlns="urn:pppa:vpk:strategija:0.1" schemaVersion="0.1" id="{DOC}">'
|
||||||
|
+ meta + abbrs + act + fund + body + annex + notes + "</PlanningDocument>\n")
|
||||||
|
# pretty-print
|
||||||
|
from lxml import etree
|
||||||
|
tree = etree.fromstring(xml.encode("utf-8"))
|
||||||
|
etree.indent(tree, space=" ")
|
||||||
|
os.makedirs(os.path.dirname(out), exist_ok=True)
|
||||||
|
with open(out, "wb") as f:
|
||||||
|
f.write(etree.tostring(tree, xml_declaration=True, encoding="UTF-8", pretty_print=True))
|
||||||
|
print("wrote", out, "actors", len(used_actors), "funding", len(used_funding), "indicator links", match_stats)
|
||||||
|
if unresolved:
|
||||||
|
print("UNRESOLVED", sorted(set(unresolved)))
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main(sys.argv[1], sys.argv[2])
|
||||||
613
tools/parse_nap.py
Normal file
613
tools/parse_nap.py
Normal file
@@ -0,0 +1,613 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""NAP2027 PDF → strukturēts JSON (starpposms; XML veido build_xml.py).
|
||||||
|
|
||||||
|
python3 tools/parse_nap.py sources/nap2027/NAP2027.pdf build/nap2027.json
|
||||||
|
|
||||||
|
Lasa vārdus ar koordinātām (pdfplumber), noņem lapu galvenes un kājenes, atdala zemsvītras piezīmes,
|
||||||
|
atpazīst nodaļas, numurētos punktus [1]–[475], indikatoru un uzdevumu tabulas (ailes pēc galvenes koordinātām)
|
||||||
|
un pielikuma pamatojuma punktus.
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
from collections import defaultdict
|
||||||
|
|
||||||
|
import pdfplumber
|
||||||
|
|
||||||
|
HEAD_TOP, FOOT_TOP = 52, 775
|
||||||
|
MARK = re.compile(r"^\[(\d{1,3})\]$")
|
||||||
|
|
||||||
|
|
||||||
|
def norm(s):
|
||||||
|
s = re.sub(r"\s+", " ", s or "").strip()
|
||||||
|
s = re.sub(r"\s+([,.;:)”])", r"\1", s)
|
||||||
|
s = re.sub(r"([“(])\s+", r"\1", s)
|
||||||
|
return repair(s)
|
||||||
|
|
||||||
|
|
||||||
|
from collections import Counter as _Counter
|
||||||
|
VOCAB = _Counter()
|
||||||
|
|
||||||
|
|
||||||
|
def load_vocab(pdf_path):
|
||||||
|
"""Word-form frequencies as pdftotext reads the body text — used to repair words split by kerning or cell hyphenation."""
|
||||||
|
import subprocess
|
||||||
|
raw = subprocess.run(["pdftotext", pdf_path, "-"], capture_output=True, text=True).stdout
|
||||||
|
VOCAB.update(w.strip(".,;:()“”\"'") for w in raw.split())
|
||||||
|
|
||||||
|
|
||||||
|
def _better(whole, *parts):
|
||||||
|
"""Join when the whole word is attested more often than any of its fragments."""
|
||||||
|
n = VOCAB.get(whole, 0) + VOCAB.get(whole.lower(), 0)
|
||||||
|
return n >= 2 and any(VOCAB.get(p, 0) < n for p in parts)
|
||||||
|
|
||||||
|
|
||||||
|
def repair(s):
|
||||||
|
if not s or not VOCAB:
|
||||||
|
return s
|
||||||
|
toks, out = s.split(" "), []
|
||||||
|
for t in toks:
|
||||||
|
if out:
|
||||||
|
a, core = out[-1], t.rstrip(".,;:)”")
|
||||||
|
ca = a.lstrip("“(")
|
||||||
|
if "/" in a or "http" in a or "www." in a:
|
||||||
|
out.append(t)
|
||||||
|
continue
|
||||||
|
# three fragments "fundamentāl a s"
|
||||||
|
if len(out) >= 2 and len(a) <= 2 and a.isalpha() and core[:1].islower() and \
|
||||||
|
_better(out[-2].lstrip("“(") + a + core, out[-2].lstrip("“("), a, core):
|
||||||
|
out.pop()
|
||||||
|
out[-1] = out[-1] + a + t
|
||||||
|
continue
|
||||||
|
# "paš - valdības"
|
||||||
|
if a == "-" and len(out) >= 2 and core[:1].islower() and _better(out[-2].lstrip("“(") + core, out[-2].lstrip("“("), core):
|
||||||
|
out.pop()
|
||||||
|
out[-1] = out[-1] + t
|
||||||
|
continue
|
||||||
|
# "aizsar- dzības"
|
||||||
|
if ca.endswith("-") and len(ca) > 2 and ca[-2].islower() and core[:1].islower() and \
|
||||||
|
(_better(ca[:-1] + core, ca[:-1], core) or VOCAB.get(core, 0) == 0):
|
||||||
|
out[-1] = a[:-1] + t
|
||||||
|
continue
|
||||||
|
# dropped capital "O glekļa", kerning "pašvaldī bas"
|
||||||
|
if ca[-1:].isalpha() and core[:1].isalpha() and core[:1].islower() and \
|
||||||
|
(_better(ca + core, ca, core) or (len(ca) == 1 and ca.isupper() and VOCAB.get((ca + core).lower(), 0) >= 2) or
|
||||||
|
(len(ca) == 1 and ca.isupper() and VOCAB.get(ca + core, 0) >= 1 and VOCAB.get(core, 0) == 0)):
|
||||||
|
out[-1] = a + t
|
||||||
|
continue
|
||||||
|
out.append(t)
|
||||||
|
return " ".join(out)
|
||||||
|
|
||||||
|
|
||||||
|
def fix_urls(s):
|
||||||
|
"""A URL broken across lines: rejoin the piece after - / _ . = ? & %."""
|
||||||
|
prev = None
|
||||||
|
while prev != s:
|
||||||
|
prev = s
|
||||||
|
s = re.sub(r"((?:https?://|www\.)\S*[-/_.%=?&])\s+(?=[\w%])", r"\1", s)
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def squash(s):
|
||||||
|
return re.sub(r"\s+", "", s or "")
|
||||||
|
|
||||||
|
|
||||||
|
def page_lines(pdf):
|
||||||
|
"""-> list of lines: {page, top, x0, words:[...], text, size, bold}; footnotes separately."""
|
||||||
|
out, notes = [], []
|
||||||
|
for pi, p in enumerate(pdf.pages):
|
||||||
|
ws = p.extract_words(extra_attrs=["size", "fontname"], keep_blank_chars=False)
|
||||||
|
ws = [w for w in ws if HEAD_TOP < w["top"] < FOOT_TOP]
|
||||||
|
# footnotes: small text at the bottom of the page
|
||||||
|
marks = [w["top"] for w in ws if w["size"] < 6.5 and w["text"].isdigit() and w["x0"] < 90 and w["top"] > 550]
|
||||||
|
fz = min(marks) - 1.5 if marks else 9999
|
||||||
|
infoot = lambda w: w["size"] < 9.5 and w["top"] >= fz
|
||||||
|
foot = [w for w in ws if infoot(w)]
|
||||||
|
body = [w for w in ws if not infoot(w) and w["size"] >= 7.5]
|
||||||
|
smalls = [w for w in ws if not infoot(w) and w["size"] < 7.5]
|
||||||
|
if foot:
|
||||||
|
fsub = [w for w in foot if w["size"] < 6.5 and w["x0"] >= 90]
|
||||||
|
foot = [dict(w) for w in foot if not (w["size"] < 6.5 and w["x0"] >= 90)]
|
||||||
|
for sw in fsub:
|
||||||
|
near = [m for m in foot if abs(m["top"] - sw["top"]) < 8 and -0.5 <= sw["x0"] - m["x1"] < 3]
|
||||||
|
if near:
|
||||||
|
m = min(near, key=lambda m: sw["x0"] - m["x1"])
|
||||||
|
m["text"] += sw["text"].translate(str.maketrans("0123456789", "₀₁₂₃₄₅₆₇₈₉"))
|
||||||
|
foot.sort(key=lambda w: (w["top"], w["x0"]))
|
||||||
|
rows_f = []
|
||||||
|
for w in foot:
|
||||||
|
if rows_f and abs(w["top"] - rows_f[-1][0]["top"]) < 2.5:
|
||||||
|
rows_f[-1].append(w)
|
||||||
|
else:
|
||||||
|
rows_f.append([w])
|
||||||
|
cur = None
|
||||||
|
for r in rows_f:
|
||||||
|
r.sort(key=lambda w: w["x0"])
|
||||||
|
if r[0]["size"] < 6.5 and r[0]["text"].isdigit() and r[0]["x0"] < 90:
|
||||||
|
cur = {"page": pi + 1, "number": int(r[0]["text"]), "text": " ".join(w["text"] for w in r[1:])}
|
||||||
|
notes.append(cur)
|
||||||
|
elif cur:
|
||||||
|
cur["text"] += " " + " ".join(w["text"] for w in r)
|
||||||
|
# merge glyphs split by kerning on one line
|
||||||
|
body.sort(key=lambda w: (round(w["top"] / 2.5), w["x0"]))
|
||||||
|
merged = []
|
||||||
|
for w in body:
|
||||||
|
if merged and abs(merged[-1]["top"] - w["top"]) < 2.5 and 0 <= w["x0"] - merged[-1]["x1"] < 0.9:
|
||||||
|
m = merged[-1]
|
||||||
|
m["text"] += w["text"]
|
||||||
|
m["x1"] = w["x1"]
|
||||||
|
m["bold"] = m["bold"] and "Bold" in w["fontname"]
|
||||||
|
else:
|
||||||
|
merged.append(dict(text=w["text"], x0=w["x0"], x1=w["x1"], top=w["top"], size=w["size"], page=pi + 1,
|
||||||
|
bold="Bold" in w["fontname"], italic="Italic" in w["fontname"]))
|
||||||
|
split = []
|
||||||
|
for w in merged:
|
||||||
|
m = re.match(r"^(\[\d{1,3}\])(.+)$", w["text"])
|
||||||
|
if m:
|
||||||
|
split.append(dict(w, text=m.group(1), x1=w["x0"] + 20))
|
||||||
|
split.append(dict(w, text=m.group(2), x0=w["x0"] + 20.5))
|
||||||
|
else:
|
||||||
|
split.append(w)
|
||||||
|
merged = split
|
||||||
|
# footnote numbers glued to a word at the same font size ("attīstībā16.")
|
||||||
|
page_notes = {n["number"] for n in notes if n["page"] == pi + 1}
|
||||||
|
for m_ in merged:
|
||||||
|
g = re.match(r"^(.*[a-zāčēģīķļņšūž”)])(\d{1,2})([.,;:]?)$", m_["text"])
|
||||||
|
if g and int(g.group(2)) in page_notes:
|
||||||
|
m_["text"] = g.group(1) + g.group(3)
|
||||||
|
m_.setdefault("refs", []).append(int(g.group(2)))
|
||||||
|
# small glyphs: subscripts (CO₂) join the word; superscript digits are footnote references
|
||||||
|
SUBS = str.maketrans("0123456789", "₀₁₂₃₄₅₆₇₈₉")
|
||||||
|
for sw in smalls:
|
||||||
|
near = [m for m in merged if abs(m["top"] - sw["top"]) < 8 and -0.5 <= sw["x0"] - m["x1"] < 3]
|
||||||
|
if not near:
|
||||||
|
continue
|
||||||
|
m = min(near, key=lambda m: sw["x0"] - m["x1"])
|
||||||
|
if sw["top"] > m["top"] + 1.5:
|
||||||
|
m["text"] += sw["text"].translate(SUBS)
|
||||||
|
m["x1"] = sw["x1"]
|
||||||
|
elif sw["text"].isdigit():
|
||||||
|
m.setdefault("refs", []).append(int(sw["text"]))
|
||||||
|
rows = defaultdict(list)
|
||||||
|
for w in merged:
|
||||||
|
rows[round(w["top"] / 2.5)].append(w)
|
||||||
|
keys = sorted(rows)
|
||||||
|
# join rows whose tops are within 2.5 pt (rounding boundary)
|
||||||
|
grouped = []
|
||||||
|
for k in keys:
|
||||||
|
if grouped and abs(rows[k][0]["top"] - grouped[-1][0]["top"]) < 2.6:
|
||||||
|
grouped[-1].extend(rows[k])
|
||||||
|
else:
|
||||||
|
grouped.append(list(rows[k]))
|
||||||
|
for g in grouped:
|
||||||
|
g.sort(key=lambda w: w["x0"])
|
||||||
|
out.append({"page": pi + 1, "top": min(w["top"] for w in g), "x0": g[0]["x0"], "words": g,
|
||||||
|
"text": " ".join(w["text"] for w in g), "size": max(w["size"] for w in g),
|
||||||
|
"bold": all(w["bold"] for w in g)})
|
||||||
|
return out, notes
|
||||||
|
|
||||||
|
|
||||||
|
SECTION_HEADS = [
|
||||||
|
("introduction", "IEVADS"), ("vision", "VĪZIJA PAR LATVIJAS NĀKOTNI 2027. GADĀ"), ("framework", "NAP2027 IETVARS"),
|
||||||
|
("strategicGoals", "NAP2027 STRATĒĢISKIE MĒRĶI"), ("spatial", "NAP2027 telpiskās attīstības perspektīva"),
|
||||||
|
("implementation", "NAP2027 īstenošanas, finansēšanas, uzraudzības un novērtēšanas process"), ("annex", "Pielikums"),
|
||||||
|
]
|
||||||
|
SUB = {
|
||||||
|
"PRIORITĀTESMĒRĶIS": "priorityGoal", "RĪCĪBASVIRZIENAMĒRĶIS": "actionLineGoal", "RĪCĪBASVIRZIENAMĒRĶI": "actionLineGoal",
|
||||||
|
"Stratēģiskomērķuindikatori": "indicators", "Rīcībasvirzienamērķaindikatori": "indicators",
|
||||||
|
"Rīcībasvirzienamērķuindikatori": "indicators", "Rīcībasvirzienauzdevumi": "tasks",
|
||||||
|
}
|
||||||
|
IND_COLS = ["no", "name", "unit", "baseYear", "baseValue", "target2024", "target2027", "source"]
|
||||||
|
TASK_COLS = ["no", "text", "responsible", "coResponsible", "funding", "indicators"]
|
||||||
|
|
||||||
|
|
||||||
|
def header_centers(lines, kind):
|
||||||
|
"""Column centres from the header words of one table (the lines between the table title and the first row)."""
|
||||||
|
ws = [w for ln in lines for w in ln["words"]]
|
||||||
|
def c(w):
|
||||||
|
return (w["x0"] + w["x1"]) / 2
|
||||||
|
def first(txt, n=0):
|
||||||
|
hits = sorted([w for w in ws if w["text"].startswith(txt)], key=lambda w: w["x0"])
|
||||||
|
return hits[n] if len(hits) > n else None
|
||||||
|
if kind == "indicators":
|
||||||
|
hs = [first("Nr"), first("Progresa") or first("Rādītājs") or first("rādītājs") or first("Indikators"), first("Mēr"), first("Bāzes", 0),
|
||||||
|
first("Bāzes", 1), first("Mērķa", 0), first("Mērķa", 1), first("Datu")]
|
||||||
|
else:
|
||||||
|
hs = [first("Nr"), first("Uzdevums"), (first("Atbildīgā") or first("Atbildī")), (first("Līdzatbildīgās") or first("Līdz")), first("Finanšu"), first("Indikators")]
|
||||||
|
if any(h is None for h in hs):
|
||||||
|
return None
|
||||||
|
return [c(h) for h in hs]
|
||||||
|
|
||||||
|
|
||||||
|
def gutters(words, centers, min_gap=2.0):
|
||||||
|
"""Column boundaries from the empty vertical strips between the words of one table page.
|
||||||
|
Between two neighbouring header centres the widest empty strip is the gutter; without one, the midpoint."""
|
||||||
|
iv = sorted((w["x0"], w["x1"]) for w in words)
|
||||||
|
gaps, end = [], None
|
||||||
|
for a, b in iv:
|
||||||
|
if end is not None and a - end >= min_gap:
|
||||||
|
gaps.append((end, a))
|
||||||
|
end = b if end is None else max(end, b)
|
||||||
|
bounds = []
|
||||||
|
for i in range(len(centers) - 1):
|
||||||
|
lo, hi = centers[i], centers[i + 1]
|
||||||
|
cand = [(g[1] - g[0], (g[0] + g[1]) / 2) for g in gaps if lo < (g[0] + g[1]) / 2 < hi]
|
||||||
|
bounds.append(max(cand)[1] if cand else (lo + hi) / 2)
|
||||||
|
return bounds
|
||||||
|
|
||||||
|
|
||||||
|
def assign(words, centers, names, bounds=None):
|
||||||
|
bounds = bounds or [(centers[i] + centers[i + 1]) / 2 for i in range(len(centers) - 1)]
|
||||||
|
cells = {n: [] for n in names}
|
||||||
|
for w in words:
|
||||||
|
cx = (w["x0"] + w["x1"]) / 2
|
||||||
|
i = sum(1 for b in bounds if cx > b)
|
||||||
|
cells[names[i]].append(w)
|
||||||
|
return cells
|
||||||
|
|
||||||
|
|
||||||
|
def cell_text(ws):
|
||||||
|
ws = sorted(ws, key=lambda w: (w.get("page", 0), round(w["top"]), w["x0"]))
|
||||||
|
t = norm(" ".join(w["text"] for w in ws))
|
||||||
|
return re.sub(r"(\d{4}/\d{1,3}) (\d)", r"\1\2", t) # "2018/2 019" broken inside a narrow cell
|
||||||
|
|
||||||
|
|
||||||
|
def split_by_gap(ws, gap=17.0):
|
||||||
|
"""Indicator names in a task row are separated by an empty line."""
|
||||||
|
ws = sorted(ws, key=lambda w: (w.get("page", 0), round(w["top"]), w["x0"]))
|
||||||
|
groups, last_top, last_page = [], None, None
|
||||||
|
for w in ws:
|
||||||
|
if last_top is None or w["top"] - last_top > gap or w.get("page") != last_page:
|
||||||
|
groups.append([])
|
||||||
|
groups[-1].append(w)
|
||||||
|
last_top, last_page = w["top"], w.get("page")
|
||||||
|
return [norm(" ".join(x["text"] for x in g)) for g in groups if g]
|
||||||
|
|
||||||
|
|
||||||
|
def table_pages(lines):
|
||||||
|
"""First pass: (kind, page) → words of table rows (≈10 pt lines between a table title and body-size text)."""
|
||||||
|
acc, mode, started, tid = defaultdict(list), None, False, 0
|
||||||
|
for ln in lines:
|
||||||
|
sq = squash(ln["text"])
|
||||||
|
if sq in SUB and ln["x0"] < 100:
|
||||||
|
mode = {"indicators": "indicator", "tasks": "task"}.get(SUB[sq])
|
||||||
|
started = False
|
||||||
|
tid += mode is not None
|
||||||
|
continue
|
||||||
|
if mode and ln["size"] >= 11:
|
||||||
|
mode = None
|
||||||
|
if mode and MARK.match(ln["words"][0]["text"]):
|
||||||
|
started = True
|
||||||
|
if mode and started and ln["size"] < 11 and not ln["text"].startswith("*"):
|
||||||
|
acc[(tid, ln["page"])] += ln["words"]
|
||||||
|
return acc
|
||||||
|
|
||||||
|
|
||||||
|
def parse(path):
|
||||||
|
pdf = pdfplumber.open(path)
|
||||||
|
load_vocab(path)
|
||||||
|
lines, notes = page_lines(pdf)
|
||||||
|
page_words = table_pages(lines)
|
||||||
|
page_bounds = {}
|
||||||
|
tid = 0
|
||||||
|
sections, items, evidence = [], [], []
|
||||||
|
sec = {"id": "front", "kind": "front", "title": "Titullapa un saīsinājumi", "parent": None}
|
||||||
|
sections.append(sec)
|
||||||
|
pr = rv = None
|
||||||
|
npr = nrv = 0
|
||||||
|
mode, centers, pending_goal = "text", None, None
|
||||||
|
centers_by = {}
|
||||||
|
area_buf = []
|
||||||
|
cur = None # current item being filled
|
||||||
|
hdr_buf = []
|
||||||
|
funding_buf = None
|
||||||
|
note_open = False
|
||||||
|
abbrs = []
|
||||||
|
problem, problem_open, problem_words = None, False, []
|
||||||
|
spatial_note, spatial_note_page = False, None
|
||||||
|
i = 0
|
||||||
|
stats = defaultdict(int)
|
||||||
|
stats_pages = []
|
||||||
|
|
||||||
|
def close():
|
||||||
|
nonlocal cur
|
||||||
|
if cur is None:
|
||||||
|
return
|
||||||
|
refs = sorted({r for w in cur.get("_words", []) for r in w.get("refs", [])})
|
||||||
|
if refs:
|
||||||
|
cur["footnoteRefs"] = refs
|
||||||
|
if cur["kind"] in ("indicator", "task"):
|
||||||
|
cols = IND_COLS if cur["kind"] == "indicator" else TASK_COLS
|
||||||
|
ws, ctr = cur.pop("_words"), cur.pop("_centers")
|
||||||
|
ws = [w for w in ws if not (w["text"] in ("(", "[") and w["x0"] < 125) and w["text"] != "["]
|
||||||
|
tid_ = cur.pop("_tid")
|
||||||
|
cells = {n: [] for n in cols}
|
||||||
|
for pg in sorted({w["page"] for w in ws}):
|
||||||
|
pw = [w for w in ws if w["page"] == pg]
|
||||||
|
key = (tid_, pg)
|
||||||
|
if key not in page_bounds:
|
||||||
|
page_bounds[key] = gutters(page_words.get(key) or pw, ctr)
|
||||||
|
b = page_bounds[key]
|
||||||
|
for k, v in assign(pw, ctr, cols, b).items():
|
||||||
|
cells[k] += v
|
||||||
|
if cur["kind"] == "indicator":
|
||||||
|
for k in cols[1:]:
|
||||||
|
cur[k] = cell_text(cells[k]) or None
|
||||||
|
else:
|
||||||
|
cur["text"] = cell_text(cells["text"])
|
||||||
|
cur["responsible"] = cell_text(cells["responsible"]) or None
|
||||||
|
cur["coResponsible"] = cell_text(cells["coResponsible"]) or None
|
||||||
|
cur["funding"] = cell_text(cells["funding"]) or None
|
||||||
|
cur["indicatorNames"] = split_by_gap(cells["indicators"])
|
||||||
|
else:
|
||||||
|
ws = cur.pop("_words")
|
||||||
|
bold = [w for w in ws if w["bold"]]
|
||||||
|
cur["text"] = norm(" ".join(w["text"] for w in ws))
|
||||||
|
cur["boldLead"] = norm(" ".join(w["text"] for w in ws[:len(ws)] if w["bold"])) if bold else None
|
||||||
|
if "_area" in cur:
|
||||||
|
cur["area"] = cur.pop("_area")[0] if cur["_area"] else None
|
||||||
|
items.append(cur)
|
||||||
|
cur = None
|
||||||
|
|
||||||
|
while i < len(lines):
|
||||||
|
ln = lines[i]
|
||||||
|
t, sq = ln["text"], squash(ln["text"])
|
||||||
|
# ---------------------------------------------------------------- big headings
|
||||||
|
if ln["size"] >= 13.5 and ln["page"] > 3:
|
||||||
|
j, title = i + 1, t
|
||||||
|
while j < len(lines) and lines[j]["size"] >= 13.5 and abs(lines[j]["size"] - ln["size"]) < 0.5 \
|
||||||
|
and lines[j]["page"] == ln["page"] and lines[j]["top"] - lines[j - 1]["top"] < 26:
|
||||||
|
title += " " + lines[j]["text"]
|
||||||
|
j += 1
|
||||||
|
title = norm(title)
|
||||||
|
close()
|
||||||
|
if funding_buf:
|
||||||
|
funding_buf = None
|
||||||
|
m = re.match(r"^Prioritāte “(.+)”$", title)
|
||||||
|
m2 = re.match(r"^Rīcības virziens “(.+)”$", title)
|
||||||
|
if sec.get("kind") == "annex" or (sections and any(s["kind"] == "annex" for s in sections)):
|
||||||
|
pass
|
||||||
|
if m:
|
||||||
|
npr += 1
|
||||||
|
pr = {"id": f"pr{npr}", "kind": "priority", "title": m.group(1), "parent": None}
|
||||||
|
sections.append(pr)
|
||||||
|
sec, rv, mode = pr, None, "text"
|
||||||
|
elif m2:
|
||||||
|
nrv += 1
|
||||||
|
rv = {"id": f"{pr['id']}.rv{nrv}", "kind": "actionLine", "title": m2.group(1), "parent": pr["id"], "funding": None}
|
||||||
|
sections.append(rv)
|
||||||
|
sec, mode = rv, "text"
|
||||||
|
elif title == "NAP2027 prioritāšu pamatojuma avoti":
|
||||||
|
mode = "annex"
|
||||||
|
else:
|
||||||
|
kind = next((k for k, h in SECTION_HEADS if squash(h) == squash(title)), None)
|
||||||
|
if kind:
|
||||||
|
sec = {"id": kind, "kind": kind, "title": title, "parent": None}
|
||||||
|
sections.append(sec)
|
||||||
|
pr = rv = None
|
||||||
|
mode = "annex" if kind == "annex" else "text"
|
||||||
|
elif ln["page"] > 4:
|
||||||
|
stats["unknown_heading"] += 1
|
||||||
|
i = j
|
||||||
|
continue
|
||||||
|
# ---------------------------------------------------------------- annex: evidence points per action line
|
||||||
|
if mode == "annex":
|
||||||
|
m = re.match(r"^Prioritāte “(.+)”$", norm(t))
|
||||||
|
m2 = re.match(r"^Rīcības virziens “(.+)”?$", norm(t))
|
||||||
|
if m:
|
||||||
|
pr = next((s for s in sections if s["kind"] == "priority" and squash(s["title"]) == squash(m.group(1))), None)
|
||||||
|
rv = None
|
||||||
|
problem, problem_open = None, False
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
if m2 and ln["x0"] < 90:
|
||||||
|
title = norm(t)
|
||||||
|
while not title.endswith("”") and i + 1 < len(lines):
|
||||||
|
i += 1
|
||||||
|
title = norm(title + " " + lines[i]["text"])
|
||||||
|
name = re.match(r"^Rīcības virziens “(.+)”$", title).group(1)
|
||||||
|
rv = next((s for s in sections if s["kind"] == "actionLine" and squash(s["title"]) == squash(name)), None)
|
||||||
|
if rv is None:
|
||||||
|
stats["annex_unknown_line"] += 1
|
||||||
|
problem, problem_open = None, False
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
m3 = re.match(r"^(\d{1,3})\.\s", t)
|
||||||
|
if m3 and ln["x0"] < 90:
|
||||||
|
problem_open = False
|
||||||
|
evidence.append({"number": int(m3.group(1)), "section": (rv or pr or {}).get("id"),
|
||||||
|
"_words": ln["words"][1:], "page": ln["page"], "problem": problem})
|
||||||
|
elif ln["bold"] and ln["x0"] < 90:
|
||||||
|
# bold problem statement that groups the following evidence points ("…:")
|
||||||
|
if problem_open and problem:
|
||||||
|
problem, problem_words = norm(problem + " " + t), problem_words + ln["words"]
|
||||||
|
else:
|
||||||
|
problem, problem_words = norm(t), list(ln["words"])
|
||||||
|
problem_open = not t.rstrip().endswith(":")
|
||||||
|
elif problem_open and ln["x0"] < 90:
|
||||||
|
# unnumbered evidence paragraph: bold lead without ":" continues as body text
|
||||||
|
evidence.append({"number": None, "section": (rv or pr or {}).get("id"), "_words": problem_words + ln["words"],
|
||||||
|
"page": ln["page"], "problem": None})
|
||||||
|
problem, problem_open, problem_words = None, False, []
|
||||||
|
elif evidence:
|
||||||
|
evidence[-1]["_words"] += ln["words"]
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
# ---------------------------------------------------------------- sub-headings
|
||||||
|
if sq in SUB and ln["x0"] < 100:
|
||||||
|
close()
|
||||||
|
kind = SUB[sq]
|
||||||
|
if kind in ("priorityGoal", "actionLineGoal"):
|
||||||
|
pending_goal, mode = kind, "text"
|
||||||
|
else:
|
||||||
|
mode, hdr_buf = kind, []
|
||||||
|
tid += 1
|
||||||
|
# header lines until the first row marker
|
||||||
|
j = i + 1
|
||||||
|
while j < len(lines) and not MARK.match(lines[j]["words"][0]["text"]):
|
||||||
|
hdr_buf.append(lines[j])
|
||||||
|
j += 1
|
||||||
|
c = header_centers(hdr_buf, kind)
|
||||||
|
if c:
|
||||||
|
centers_by[kind] = c
|
||||||
|
else:
|
||||||
|
stats["header_fallback_" + kind] += 1
|
||||||
|
stats.setdefault("header_fallback_pages", []).append(ln["page"])
|
||||||
|
centers = centers_by.get(kind)
|
||||||
|
i = j
|
||||||
|
continue
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
# ---------------------------------------------------------------- indicative funding of an action line
|
||||||
|
if sq.startswith("Rīcībasvirzienapasākumu") or funding_buf is not None:
|
||||||
|
close()
|
||||||
|
funding_buf = (funding_buf or "") + " " + t
|
||||||
|
if "EUR" in t:
|
||||||
|
m = re.search(r"apjoms\s+([\d\s,]+)\s*milj\.\s*EUR", norm(funding_buf))
|
||||||
|
if rv is not None:
|
||||||
|
rv["funding"] = {"text": norm(funding_buf), "millionEur": m.group(1).replace(" ", "") if m else None}
|
||||||
|
funding_buf = None
|
||||||
|
mode = "text"
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
# ---------------------------------------------------------------- spatial perspective: area | items
|
||||||
|
if sec.get("kind") == "spatial" and (t.startswith("Pilsētu un lauku mijiedarbības dimensijas") or spatial_note):
|
||||||
|
if MARK.search(t) or ln["page"] != spatial_note_page and spatial_note:
|
||||||
|
spatial_note = False
|
||||||
|
else:
|
||||||
|
close()
|
||||||
|
if not spatial_note:
|
||||||
|
sec.setdefault("notes", []).append({"area": area_buf[0] if area_buf else None, "text": norm(t)})
|
||||||
|
spatial_note, spatial_note_page = True, ln["page"]
|
||||||
|
else:
|
||||||
|
sec["notes"][-1]["text"] = norm(sec["notes"][-1]["text"] + ("\n" if t.startswith(("–", "−")) else " ") + t)
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
if sec.get("kind") == "spatial":
|
||||||
|
mk_i = next((k for k, w in enumerate(ln["words"]) if MARK.match(w["text"])), None)
|
||||||
|
if mk_i is not None and ln["words"][mk_i]["x0"] > 180:
|
||||||
|
close()
|
||||||
|
left = [w for w in ln["words"][:mk_i] if w["x0"] < 195]
|
||||||
|
if left:
|
||||||
|
area_buf[:] = [norm(" ".join(w["text"] for w in left))]
|
||||||
|
cur = {"number": int(MARK.match(ln["words"][mk_i]["text"]).group(1)), "kind": "spatialDirection",
|
||||||
|
"section": "spatial", "page": ln["page"], "_words": ln["words"][mk_i + 1:], "_area": area_buf}
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
if cur is not None and cur.get("_area") is not None:
|
||||||
|
left = [w for w in ln["words"] if w["x0"] < 195]
|
||||||
|
right = [w for w in ln["words"] if w["x0"] >= 195]
|
||||||
|
if left and not right and len(left) <= 5:
|
||||||
|
area_buf[0] = norm(area_buf[0] + " " + " ".join(w["text"] for w in left)) if area_buf else norm(" ".join(w["text"] for w in left))
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
if left and right and right[0]["x0"] > 190 and len(left) <= 4:
|
||||||
|
area_buf[0] = norm(area_buf[0] + " " + " ".join(w["text"] for w in left))
|
||||||
|
cur["_words"] += right
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
# ---------------------------------------------------------------- table notes ("* ...", 9 pt)
|
||||||
|
if (t.startswith("*") and ln["size"] < 9.5) or (note_open and ln["size"] < 9.5 and not MARK.match(ln["words"][0]["text"])):
|
||||||
|
if t.startswith("*"):
|
||||||
|
close()
|
||||||
|
sec.setdefault("notes", []).append(norm(t.lstrip("* ")))
|
||||||
|
else:
|
||||||
|
sec["notes"][-1] = norm(sec["notes"][-1] + " " + t)
|
||||||
|
note_open = True
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
note_open = False
|
||||||
|
# ---------------------------------------------------------------- thematic sub-sections inside an action line
|
||||||
|
if rv is not None and 11.5 < ln["size"] < 13.5 and t.startswith("“") and norm(t).endswith("”") and ln["x0"] < 90:
|
||||||
|
close()
|
||||||
|
n_th = sum(1 for x in sections if x.get("parent") == rv["id"]) + 1
|
||||||
|
sec = {"id": f"{rv['id']}.t{n_th}", "kind": "theme", "title": norm(t).strip("“”").strip(), "parent": rv["id"]}
|
||||||
|
sections.append(sec)
|
||||||
|
mode = "text"
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
# ---------------------------------------------------------------- numbered items
|
||||||
|
first = ln["words"][0]
|
||||||
|
mk = MARK.match(first["text"])
|
||||||
|
in_table = mode in ("indicators", "tasks")
|
||||||
|
if in_table and not mk and ln["size"] >= 11:
|
||||||
|
# body-size text ends the table (table cells are ≈10 pt)
|
||||||
|
if True:
|
||||||
|
close()
|
||||||
|
mode = "text"
|
||||||
|
in_table = False
|
||||||
|
if mk:
|
||||||
|
close()
|
||||||
|
n = int(mk.group(1))
|
||||||
|
if in_table and ln["size"] < 11:
|
||||||
|
kind = "indicator" if mode == "indicators" else "task"
|
||||||
|
cur = {"number": n, "kind": kind, "section": (sec if sec.get("kind") == "theme" else (rv or pr or sec))["id"], "page": ln["page"],
|
||||||
|
"_words": ln["words"][1:], "_centers": centers, "_tid": tid}
|
||||||
|
else:
|
||||||
|
if in_table:
|
||||||
|
mode = "text"
|
||||||
|
k = "text"
|
||||||
|
if pending_goal:
|
||||||
|
k = pending_goal
|
||||||
|
cur = {"number": n, "kind": k, "section": (sec if sec.get("kind") == "theme" else (rv or pr or sec))["id"], "page": ln["page"], "_words": ln["words"][1:]}
|
||||||
|
if pending_goal:
|
||||||
|
pending_goal = None if pending_goal == "priorityGoal" else pending_goal
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
if cur is not None:
|
||||||
|
cur["_words"] += [w for w in ln["words"] if not (w["text"] == "[" )]
|
||||||
|
elif sec["kind"] == "front" and ln["page"] in (3, 4) and t != "IZMANTOTIE SAĪSINĀJUMI":
|
||||||
|
m = re.match(r"^(.+?)\s+–\s+(.+)$", t)
|
||||||
|
if m and ln["x0"] < 90 and len(m.group(1)) <= 40:
|
||||||
|
abbrs.append({"abbr": m.group(1).strip(), "meaning": m.group(2).strip()})
|
||||||
|
elif abbrs:
|
||||||
|
abbrs[-1]["meaning"] = norm(abbrs[-1]["meaning"] + " " + t)
|
||||||
|
elif sec["kind"] == "front":
|
||||||
|
pass # title page and table of contents
|
||||||
|
else:
|
||||||
|
stats["orphan_line"] += 1
|
||||||
|
stats.setdefault("orphans", []).append([ln["page"], (sec or {}).get("id"), " ".join(w["text"] for w in ln["words"])[:110]])
|
||||||
|
i += 1
|
||||||
|
close()
|
||||||
|
|
||||||
|
# goals: after RĪCĪBAS VIRZIENA MĒRĶIS(-I) only bold items are goals; the first non-bold item is context
|
||||||
|
for it in items:
|
||||||
|
if it["kind"] == "actionLineGoal" and not (it.get("boldLead") and len(it["boldLead"]) >= 0.6 * len(it["text"])):
|
||||||
|
it["kind"] = "text"
|
||||||
|
# once a non-goal follows, later items in the section are context
|
||||||
|
seen_text = set()
|
||||||
|
for it in items:
|
||||||
|
if it["kind"] == "text":
|
||||||
|
seen_text.add(it["section"])
|
||||||
|
elif it["kind"] == "actionLineGoal" and it["section"] in seen_text:
|
||||||
|
it["kind"] = "text"
|
||||||
|
# strategic goals: paragraphs with a bold lead in the strategic goals section
|
||||||
|
for it in items:
|
||||||
|
if it["section"] == "strategicGoals" and it["kind"] == "text" and it.get("boldLead"):
|
||||||
|
it["kind"] = "strategicGoal"
|
||||||
|
it["title"] = it["boldLead"]
|
||||||
|
for e in evidence:
|
||||||
|
refs = sorted({r for w in e["_words"] for r in w.get("refs", [])})
|
||||||
|
if refs:
|
||||||
|
e["footnoteRefs"] = refs
|
||||||
|
e["text"] = norm(" ".join(w["text"] for w in e.pop("_words")))
|
||||||
|
urls = re.findall(r"(?:https?://|www\.)\S+", re.sub(r"(?<=[-/_.%=?&])\s+(?=\S)", "", e["text"]))
|
||||||
|
clean = []
|
||||||
|
for u in urls:
|
||||||
|
u = u.rstrip(".,;")
|
||||||
|
while u.endswith(")") and u.count(")") > u.count("("):
|
||||||
|
u = u[:-1]
|
||||||
|
clean.append(u.rstrip(".,;"))
|
||||||
|
e["urls"] = clean
|
||||||
|
return {"abbreviations": abbrs, "sections": sections, "items": items, "evidence": evidence, "footnotes": [dict(n, text=fix_urls(norm(n["text"]))) for n in notes], "stats": dict(stats)}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
doc = parse(sys.argv[1])
|
||||||
|
out = sys.argv[2]
|
||||||
|
with open(out, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(doc, f, ensure_ascii=False, indent=1)
|
||||||
|
from collections import Counter
|
||||||
|
print("sections", Counter(s["kind"] for s in doc["sections"]))
|
||||||
|
print("items", len(doc["items"]), Counter(i["kind"] for i in doc["items"]))
|
||||||
|
print("evidence", len(doc["evidence"]), "footnotes", len(doc["footnotes"]), "stats", doc["stats"])
|
||||||
69
tools/validate.py
Normal file
69
tools/validate.py
Normal file
@@ -0,0 +1,69 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Strategy-as-Code pārbaudītājs.
|
||||||
|
|
||||||
|
python3 tools/validate.py [--registry organizacijas.xml]
|
||||||
|
|
||||||
|
Katram dokumentam katalogā planosanas-dokumenti.yaml:
|
||||||
|
- datne atbilst XSD shēmai schemas/strategija-0.1.xsd (arī atslēgas un atsauces shēmā);
|
||||||
|
- datnes nosaukums = dokumenta identifikators;
|
||||||
|
- avota PDF sha256 sakrīt ar Metadata/Source/SHA256;
|
||||||
|
- visi VPK ID (Actors, Responsible, CoResponsible, Abbreviation, Developer) ir Valsts institūciju reģistrā;
|
||||||
|
- pārbaudes atskaite verification/<id>.verify.json ir par šo pašu avota PDF un ir izturēta.
|
||||||
|
Reģistrs pēc noklusējuma — ProcessGit Valdibas-Deklaracija-as-Code data/organizacijas.xml.
|
||||||
|
"""
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
import yaml
|
||||||
|
from lxml import etree
|
||||||
|
|
||||||
|
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
NS = {"s": "urn:pppa:vpk:strategija:0.1"}
|
||||||
|
REG_URL = "https://processgit.org/Valsts-Pirmkods/Valdibas-Deklaracija-as-Code/raw/branch/main/data/organizacijas.xml"
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
reg_src = sys.argv[sys.argv.index("--registry") + 1] if "--registry" in sys.argv else REG_URL
|
||||||
|
reg = etree.parse(reg_src if os.path.exists(reg_src) else urllib.request.urlopen(reg_src))
|
||||||
|
ids = {o.get("id") for o in reg.getroot() if etree.QName(o).localname == "Organization"}
|
||||||
|
xsd = etree.XMLSchema(etree.parse(os.path.join(ROOT, "schemas/strategija-0.1.xsd")))
|
||||||
|
cat = yaml.safe_load(open(os.path.join(ROOT, "planosanas-dokumenti.yaml"), encoding="utf-8"))
|
||||||
|
errors, n = [], 0
|
||||||
|
for d in cat["documents"]:
|
||||||
|
if not d.get("as_code", {}).get("path"):
|
||||||
|
continue
|
||||||
|
n += 1
|
||||||
|
path = os.path.join(ROOT, d["as_code"]["path"])
|
||||||
|
doc = etree.parse(path)
|
||||||
|
if not xsd.validate(doc):
|
||||||
|
errors += [f"{d['id']}: XSD {e.line}: {e.message}" for e in list(xsd.error_log)[:10]]
|
||||||
|
root = doc.getroot()
|
||||||
|
if root.get("id") != d["id"] or os.path.basename(path) != d["id"] + ".xml":
|
||||||
|
errors.append(f"{d['id']}: datnes nosaukums vai id neatbilst katalogam")
|
||||||
|
src = os.path.join(ROOT, root.findtext("s:Metadata/s:Source/s:File", namespaces=NS))
|
||||||
|
sha = hashlib.sha256(open(src, "rb").read()).hexdigest()
|
||||||
|
if sha != root.findtext("s:Metadata/s:Source/s:SHA256", namespaces=NS):
|
||||||
|
errors.append(f"{d['id']}: avota sha256 nesakrīt")
|
||||||
|
used = {e.get("org") for e in root.iter() if e.get("org")}
|
||||||
|
missing = sorted(used - ids)
|
||||||
|
if missing:
|
||||||
|
errors.append(f"{d['id']}: VPK ID nav reģistrā: {', '.join(missing)}")
|
||||||
|
vpath = os.path.join(ROOT, "verification", d["id"] + ".verify.json")
|
||||||
|
if os.path.exists(vpath):
|
||||||
|
v = json.load(open(vpath, encoding="utf-8"))
|
||||||
|
if v.get("pdf_sha256") != sha or not v.get("passed"):
|
||||||
|
errors.append(f"{d['id']}: pārbaudes atskaite nav izturēta vai ir par citu avota datni")
|
||||||
|
else:
|
||||||
|
errors.append(f"{d['id']}: nav pārbaudes atskaites {vpath}")
|
||||||
|
print(f"{d['id']}: {len(doc.findall('.//s:Item', NS))} punkti, {len(used)} VPK ID")
|
||||||
|
for e in errors:
|
||||||
|
print("KĻŪDA", e)
|
||||||
|
print("pārbaudīti dokumenti:", n, "kļūdas:", len(errors))
|
||||||
|
sys.exit(1 if errors else 0)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
134
tools/verify_nap2027.py
Normal file
134
tools/verify_nap2027.py
Normal file
@@ -0,0 +1,134 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""NAP2027 datnes pārbaude pret avota PDF ar citu PDF nolasītāju (poppler pdftotext), nevis to, ar kuru datne veidota (pdfplumber).
|
||||||
|
|
||||||
|
python3 tools/verify_nap2027.py data/nap2027/lv-nap2027.xml sources/nap2027/NAP2027.pdf verification/lv-nap2027.verify.json
|
||||||
|
|
||||||
|
Pārbauda:
|
||||||
|
1. numurētie punkti [1]–[N] — katrs PDF drukātais numurs ir datnē tieši vienreiz, un otrādi;
|
||||||
|
2. teksta punkti — punkta burtu secība ir PDF tekstā (bez atstarpēm, pieturzīmēm un cipariem; lapu galvenes izņemtas);
|
||||||
|
3. tabulu rindas (indikatori, uzdevumi) — katrs datnes vārds atrodams PDF teksta fragmentā starp [n] un [n+1]
|
||||||
|
(vārds vai tā daļas, ja PDF to pārnesis citā rindā);
|
||||||
|
4. pamatojuma punkti — burtu secība ir PDF tekstā;
|
||||||
|
5. uzdevumi — katram ir atbildīgā institūcija ar VPK ID; indikatori — nosaukums un vismaz viena vērtība.
|
||||||
|
"""
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
|
||||||
|
from lxml import etree
|
||||||
|
|
||||||
|
NS = {"s": "urn:pppa:vpk:strategija:0.1"}
|
||||||
|
HEADER = "latvijasnacionālaisattīstībasplānsgadam"
|
||||||
|
SUBS = str.maketrans("₀₁₂₃₄₅₆₇₈₉", "0123456789")
|
||||||
|
|
||||||
|
|
||||||
|
def letters(s):
|
||||||
|
return re.sub(r"[^a-zāčēģīķļņšūž]", "", (s or "").lower())
|
||||||
|
|
||||||
|
|
||||||
|
def words(s):
|
||||||
|
return [w for w in re.findall(r"[0-9a-zāčēģīķļņšūž]+", (s or "").lower().translate(SUBS))]
|
||||||
|
|
||||||
|
|
||||||
|
def runs(t, flat, min_run=8):
|
||||||
|
"""Number of contiguous pieces of t found in flat (PDF reading order may interleave footnotes or table cells);
|
||||||
|
None if some piece shorter than min_run is not found."""
|
||||||
|
n, i = 0, 0
|
||||||
|
while i < len(t):
|
||||||
|
lo, hi = 0, len(t) - i
|
||||||
|
while lo < hi:
|
||||||
|
m = (lo + hi + 1) // 2
|
||||||
|
if t[i:i + m] in flat:
|
||||||
|
lo = m
|
||||||
|
else:
|
||||||
|
hi = m - 1
|
||||||
|
if lo < min(min_run, len(t) - i):
|
||||||
|
return None
|
||||||
|
n, i = n + 1, i + lo
|
||||||
|
return n
|
||||||
|
|
||||||
|
|
||||||
|
def main(xml_path, pdf_path, out):
|
||||||
|
raw = subprocess.run(["pdftotext", pdf_path, "-"], capture_output=True, text=True, check=True).stdout
|
||||||
|
lay = subprocess.run(["pdftotext", "-layout", pdf_path, "-"], capture_output=True, text=True, check=True).stdout
|
||||||
|
flat = letters(raw).replace(HEADER, "")
|
||||||
|
doc = etree.parse(xml_path)
|
||||||
|
items = doc.findall(".//s:Item", NS)
|
||||||
|
res = {"file": xml_path, "pdf_sha256": hashlib.sha256(open(pdf_path, "rb").read()).hexdigest(), "checks": {}}
|
||||||
|
|
||||||
|
# 1. numbering
|
||||||
|
printed = sorted({int(m) for m in re.findall(r"\[(\d{1,3})\]", lay)})
|
||||||
|
have = [int(i.get("n")) for i in items]
|
||||||
|
dup = sorted({n for n in have if have.count(n) > 1})
|
||||||
|
res["checks"]["numbering"] = {"printed": len(printed), "in_file": len(have), "missing": sorted(set(printed) - set(have)),
|
||||||
|
"extra": sorted(set(have) - set(printed)), "duplicates": dup}
|
||||||
|
|
||||||
|
# layout blocks between consecutive markers
|
||||||
|
pos = {int(m.group(1)): m.end() for m in re.finditer(r"\[(\d{1,3})\]", lay)}
|
||||||
|
order = sorted(pos)
|
||||||
|
|
||||||
|
def block(n):
|
||||||
|
k = order.index(n)
|
||||||
|
end = pos[order[k + 1]] if k + 1 < len(order) else len(lay)
|
||||||
|
return lay[pos[n]:end]
|
||||||
|
|
||||||
|
text_bad, text_split, row_bad, row_words, text_n, row_n = [], [], [], 0, 0, 0
|
||||||
|
for it in items:
|
||||||
|
n, kind = int(it.get("n")), it.get("kind")
|
||||||
|
if kind in ("indicator", "task"):
|
||||||
|
row_n += 1
|
||||||
|
toks = set(words(block(n)))
|
||||||
|
# footnote numbers printed after a word or number ("piesaisti10", "4977"): also try without them
|
||||||
|
toks |= {re.sub(r"\d{1,2}$", "", w) for w in toks if re.search(r"[a-zāčēģīķļņšūž]\d{1,2}$", w)}
|
||||||
|
toks |= {w[:-k] for w in toks if w.isdigit() and len(w) > 2 for k in (1, 2)}
|
||||||
|
vals = [x.text for x in it.iter() if x.text and x.text.strip() and etree.QName(x).localname not in ("Item",)]
|
||||||
|
vals += [a.get("printed") for a in it.iter() if a.get("printed")]
|
||||||
|
missing = []
|
||||||
|
for w in words(" ".join(vals)):
|
||||||
|
row_words += 1
|
||||||
|
if w in toks:
|
||||||
|
continue
|
||||||
|
if any(w[:k] in toks and w[k:] in toks for k in range(1, len(w))):
|
||||||
|
continue
|
||||||
|
if any(w[:k] in toks and w[k:j] in toks and w[j:] in toks for k in range(1, len(w)) for j in range(k + 1, len(w))):
|
||||||
|
continue
|
||||||
|
missing.append(w)
|
||||||
|
if missing:
|
||||||
|
row_bad.append({"n": n, "missing_words": missing[:10]})
|
||||||
|
else:
|
||||||
|
text_n += 1
|
||||||
|
ks = [runs(letters(x.text), flat) for x in it.findall("s:Text", NS) + it.findall("s:Area", NS)]
|
||||||
|
k = None if None in ks else max(ks)
|
||||||
|
if k is None or k > 4:
|
||||||
|
text_bad.append(n)
|
||||||
|
elif k > 1:
|
||||||
|
text_split.append(n)
|
||||||
|
res["checks"]["text_items"] = {"checked": text_n, "not_found_verbatim": text_bad,
|
||||||
|
"found_in_2_to_4_pieces": text_split}
|
||||||
|
res["checks"]["table_rows"] = {"checked": row_n, "words": row_words, "rows_with_missing_words": row_bad}
|
||||||
|
|
||||||
|
ev = doc.findall(".//s:Evidence", NS)
|
||||||
|
ev_bad = [e.get("id") for e in ev if (runs(letters((e.findtext("s:Problem", namespaces=NS) or "") + e.findtext("s:Text", namespaces=NS)), flat) or 99) > 4]
|
||||||
|
res["checks"]["evidence"] = {"checked": len(ev), "numbers_continuous": [int(e.get("n")) for e in ev if e.get("n")] == list(range(1, len([e for e in ev if e.get("n")]) + 1)),
|
||||||
|
"not_found_verbatim": ev_bad}
|
||||||
|
|
||||||
|
tasks = doc.findall(".//s:Task", NS)
|
||||||
|
no_resp = [t.getparent().get("n") for t in tasks if not t.findall("s:Responsible/s:Actor[@org]", NS)]
|
||||||
|
inds = doc.findall(".//s:Indicator", NS)
|
||||||
|
no_val = [i.getparent().get("n") for i in inds if not (i.findtext("s:BaseValue", namespaces=NS) or i.findall("s:Target", NS))]
|
||||||
|
res["checks"]["completeness"] = {"tasks": len(tasks), "tasks_without_responsible_vpk_id": no_resp,
|
||||||
|
"indicators": len(inds), "indicators_without_values": no_val}
|
||||||
|
c = res["checks"]
|
||||||
|
res["passed"] = not (c["numbering"]["missing"] or c["numbering"]["extra"] or c["numbering"]["duplicates"] or text_bad
|
||||||
|
or row_bad or ev_bad or no_resp)
|
||||||
|
with open(out, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(res, f, ensure_ascii=False, indent=1)
|
||||||
|
print(json.dumps({k: {kk: (vv if not isinstance(vv, list) else (len(vv) if len(vv) > 12 else vv)) for kk, vv in v.items()}
|
||||||
|
for k, v in c.items()}, ensure_ascii=False, indent=1))
|
||||||
|
print("passed", res["passed"])
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main(*sys.argv[1:4])
|
||||||
44
verification/lv-nap2027.verify.json
Normal file
44
verification/lv-nap2027.verify.json
Normal file
@@ -0,0 +1,44 @@
|
|||||||
|
{
|
||||||
|
"file": "data/nap2027/lv-nap2027.xml",
|
||||||
|
"pdf_sha256": "3da6994517bf43d42482dde00f46d4f30776f754830dfccb085ad8def2c9cd56",
|
||||||
|
"checks": {
|
||||||
|
"numbering": {
|
||||||
|
"printed": 475,
|
||||||
|
"in_file": 475,
|
||||||
|
"missing": [],
|
||||||
|
"extra": [],
|
||||||
|
"duplicates": []
|
||||||
|
},
|
||||||
|
"text_items": {
|
||||||
|
"checked": 220,
|
||||||
|
"not_found_verbatim": [],
|
||||||
|
"found_in_2_to_4_pieces": [
|
||||||
|
57,
|
||||||
|
255,
|
||||||
|
445,
|
||||||
|
452,
|
||||||
|
453,
|
||||||
|
454,
|
||||||
|
455,
|
||||||
|
457
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"table_rows": {
|
||||||
|
"checked": 255,
|
||||||
|
"words": 9684,
|
||||||
|
"rows_with_missing_words": []
|
||||||
|
},
|
||||||
|
"evidence": {
|
||||||
|
"checked": 154,
|
||||||
|
"numbers_continuous": true,
|
||||||
|
"not_found_verbatim": []
|
||||||
|
},
|
||||||
|
"completeness": {
|
||||||
|
"tasks": 124,
|
||||||
|
"tasks_without_responsible_vpk_id": [],
|
||||||
|
"indicators": 131,
|
||||||
|
"indicators_without_values": []
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"passed": true
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user