Official AI Store

Datasets for
your AI models

We are an AI lab and curate datasets for every training phase: from pretraining to supervised fine‑tuning.

Datasets

High‑quality data to train models

Proprietary datasets developed by the Egomnia team. Clean, structured, ready to use.

Progetto Talia dataset cover
Pretraining

Progetto Talia

Raw Italian text corpus. Over 360 contents from Progetto Talia, the personal and professional growth platform founded by Matteo Achilli. Authentic material for language model pretraining.

Download a free sample →
Format
.txt (raw text)
Size
774 Kb
Language
Italian
Category
Personal, professional and personal finance growth content
Commercial AI Dataset License

Purchasers may use the dataset to train, fine‑tune and evaluate AI and machine learning models and may freely use the resulting models, including commercially, such as services, APIs or products based on them. Resale, sublicensing or redistribution of the dataset or substantial parts of it to third parties is not permitted. It is mandatory to credit Egomnia S.p.A. as the source of the dataset in any technical documentation, model card or public description of the system developed.

Userprompt.ai dataset cover
Pretraining

Userprompt.ai

Raw text corpus in the AI domain taken from the AI education course created by Matteo Achilli. Content to share AI knowledge with a wide audience.

Download a free sample →
Format
.txt (raw text)
Size
78 Kb
Language
Italian
Category
Artificial Intelligence, Machine Learning and Deep Learning
Commercial AI Dataset License

Purchasers may use the dataset to train, fine‑tune and evaluate AI and machine learning models and may freely use the resulting models, including commercially, such as services, APIs or products based on them. Resale, sublicensing or redistribution of the dataset or substantial parts of it to third parties is not permitted. It is mandatory to credit Egomnia S.p.A. as the source of the dataset in any technical documentation, model card or public description of the system developed.

Progetto Italia dataset cover
Pretraining

Progetto Italia

Raw Italian text corpus dedicated to promoting Italian excellence and culture. A project curated by Matteo Achilli to highlight the cultural, artistic and entrepreneurial heritage of our country.

Download a free sample →
Format
.txt (raw text)
Size
25 Kb
Language
Italian
Category
Italian culture, curiosities & excellence
Commercial AI Dataset License

Purchasers may use the dataset to train, fine‑tune and evaluate AI and machine learning models and may freely use the resulting models, including commercially, such as services, APIs or products based on them. Resale, sublicensing or redistribution of the dataset or substantial parts of it to third parties is not permitted. It is mandatory to credit Egomnia S.p.A. as the source of the dataset in any technical documentation, model card or public description of the system developed.

Egomnia.com dataset cover
Pretraining

Egomnia.com

Raw text corpus composed of over 100,000 cover letters generated by frontier AI, created using anonymized data from Egomnia.com data entry, the first Italian social network dedicated to matching companies and candidates, launched in 2012.

Format
.txt (raw text)
Size
31 MB
Language
Italian
Category
Human Resources
Commercial AI Dataset License

Purchasers may use the dataset to train, fine‑tune and evaluate AI and machine learning models and may freely use the resulting models, including commercially, such as services, APIs or products based on them. Resale, sublicensing or redistribution of the dataset or substantial parts of it to third parties is not permitted. It is mandatory to credit Egomnia S.p.A. as the source of the dataset in any technical documentation, model card or public description of the system developed.

EmmaSFT 6.0 dataset cover
Supervised Fine‑Tuning

EmmaSFT 6.0

Dataset of over 27 thousand prompt‑response pairs in .json format used for partial fine‑tuning of Emma 6, Egomnia’s open‑source model. Designed for conversational capabilities and instruction alignment.

Download a free sample →
Format
.json (prompt‑response)
Target model
Emma 6
Language
Italian
Category
Conversational fine‑tuning
Commercial AI Dataset License

Purchasers may use the dataset to train, fine‑tune and evaluate AI and machine learning models and may freely use the resulting models, including commercially, such as services, APIs or products based on them. Resale, sublicensing or redistribution of the dataset or substantial parts of it to third parties is not permitted. It is mandatory to credit Egomnia S.p.A. as the source of the dataset in any technical documentation, model card or public description of the system developed.

Academic partnership

Are you a university?

We believe in research. All datasets are free for accredited universities and research institutions. Faculty can request full access for teaching and scientific research.

Contact us for dataset access

«We want Egomnia datasets to be used in academic research projects across Italy.»