diff --git a/pages/interns/dries_huybens/README.md b/pages/interns/dries_huybens/README.md index d0449fd2..a3fde815 100644 --- a/pages/interns/dries_huybens/README.md +++ b/pages/interns/dries_huybens/README.md @@ -42,7 +42,7 @@ Project tasks update: - Prepare a multilingual parallel corpus from OPUS data - Prepare dataset for visual language model evaluation from foto - Prepare dataset for visual language model evaluation from wikipedia. You can use https://huggingface.co/datasets/wikimedia/wit_base . -- Prepare corpus of text and image data from pravda.sk . +- Prepare corpus of text and image data from pravda.sk . For each subdomain, create - raw HTML data with images. Then create corpus of images plus descriptions, subtitles and tags. Then create corpus of extracted text. You can use docling , trafilatura for text extration. Or design your own parser. Give source codes to a git repository with documentation. Put data files to school server. Do not download too fast, so our IP does not reveive a ban. Outupus: