PROJECTS


DATASETS

LAION-400M

image/text

Status: Released


Formerly known as crawling@home (C@H), an openly accessible 400M image-text-pair dataset.

LAION5B

image/text

Status: Released


A dataset consisting of 5.85 billion CLIP-filtered image-text pairs, featuring several nearest neighbor indices, an improved web-interface for exploration and subset generation, and detection scores for watermark, NSFW, and toxic content detection.

Laion-coco

image/text

Status: Released


600M captions generated using BLIP from Laion2B-en.

Laion translated

image/text

Status: Released


3B translated samples from Laion5B.

Clip H/14

image/text

Status: Released


The largest open source clip.

LAION5B High-Res

image/text

Status: Released


A subset of the LAION5B database, with high resolution images over 1024x1024, containing 170 million samples.

LAION Aesthetics

image/text

Status: Released


A subset of LAION5B that has been estimated by a model trained on top of clip embeddings to contain only aestheticly pleasing images.

LAION-3D

3d/image/text

Status: Started


An effort to create a large-scale dataset consisting of 3D models and descriptor pairs.

Audio Dataset

text/audio

Status: Started


An audio dataset for training CLAP and other models, containing a raw and processed dataset, the latter containing .flac files with captions, labels, and other metadata.

Watermark Detection

image/text

Contrastive

Status: Released


A repository containing datasets to train a watermark classifier.

LAION-BVD

video/text

Status: Released


A large-scale open video dataset for multimodal learning, containing 1.3B platform-specific video URLs collected from CommonCrawl, 80M downloaded videos with a total duration of 10 million hours, 55 million captioned video clips, and 300 million captioned video frames.

MODELS