zz
This commit is contained in:
parent
c438574986
commit
45a4828511
@ -42,7 +42,7 @@ Project tasks update:
|
||||
- Prepare a multilingual parallel corpus from OPUS data
|
||||
- Prepare dataset for visual language model evaluation from foto
|
||||
- Prepare dataset for visual language model evaluation from wikipedia. You can use https://huggingface.co/datasets/wikimedia/wit_base .
|
||||
- Prepare corpus of text and image data from pravda.sk .
|
||||
- Prepare corpus of text and image data from pravda.sk . For each subdomain, create - raw HTML data with images. Then create corpus of images plus descriptions, subtitles and tags. Then create corpus of extracted text. You can use docling , trafilatura for text extration. Or design your own parser. Give source codes to a git repository with documentation. Put data files to school server. Do not download too fast, so our IP does not reveive a ban.
|
||||
|
||||
Outupus:
|
||||
|
||||
|
||||
Loading…
Reference in New Issue
Block a user