This commit is contained in:
Daniel Hládek 2026-08-13 14:10:07 +02:00
parent c438574986
commit 45a4828511

View File

@ -42,7 +42,7 @@ Project tasks update:
- Prepare a multilingual parallel corpus from OPUS data
- Prepare dataset for visual language model evaluation from foto
- Prepare dataset for visual language model evaluation from wikipedia. You can use https://huggingface.co/datasets/wikimedia/wit_base .
- Prepare corpus of text and image data from pravda.sk .
- Prepare corpus of text and image data from pravda.sk . For each subdomain, create - raw HTML data with images. Then create corpus of images plus descriptions, subtitles and tags. Then create corpus of extracted text. You can use docling , trafilatura for text extration. Or design your own parser. Give source codes to a git repository with documentation. Put data files to school server. Do not download too fast, so our IP does not reveive a ban.
Outupus: