dmytro_ushatenko/pages/interns/dries_huybens/README.md
2026-08-13 14:10:07 +02:00

52 lines
2.0 KiB
Markdown

---
title: Dries Huybens
published: true
taxonomy:
category: [erasmus]
tag: [nlp, ie, rag, medical]
author: Daniel Hladek
---
ERASMUS Intern Summer 2026, 20 July - 30 August
Topic:
Multilingual Knowledge Graphs from medical data
Goal:
- Construct a knowledge graph from medical package inserts in multiple languages
- Utilize the graph in an intelligent agent that recommends medication.
Tasks:
- Continue project of [Bogdan Paul Chis](/interns/bogdan_paul_chis)
- Study repositories:
- https://github.com/chis-facultate/erasmus-kosice
- https://github.com/hladek/mul-me-kg
- https://github.com/hkuds/lightrag
- Learn intelligent agents and generative models - OpenAI API, Agent frameworks, RAG systems.
- Learn about knowledge graphs and GraphRAG. Read several research papers.
- Prepare a Python based workflow, use git code repository
- Visualize the graph
- Prepare an agent that utilizes the unstructured data and graph-data.
- Evaluate the agent using DeepEval or RAGAS.
- Write a report
- Put all code to GIT
Project tasks update:
- Prepare a multilingual parallel corpus from OPUS data
- Prepare dataset for visual language model evaluation from foto
- Prepare dataset for visual language model evaluation from wikipedia. You can use https://huggingface.co/datasets/wikimedia/wit_base .
- Prepare corpus of text and image data from pravda.sk . For each subdomain, create - raw HTML data with images. Then create corpus of images plus descriptions, subtitles and tags. Then create corpus of extracted text. You can use docling , trafilatura for text extration. Or design your own parser. Give source codes to a git repository with documentation. Put data files to school server. Do not download too fast, so our IP does not reveive a ban.
Outupus:
- multilingual corpus was analyzed with embedding models. Semantic overlap between languages is low, so multilngual parallel corpus will be small.
- A [corpus](https://huggingface.co/datasets/driesaster/fotkyzadarmo_slovak) from FotkyZadarmo.sk