Explore 3 tools and 26 datasets tagged with Machine Learning for Python data engineering.
Machine learning datasets and tools are used to train, validate, and deploy predictive models in Python data engineering workflows. This tag covers training datasets, feature stores, ML pipeline orchestration, and model serving tools. Python libraries like scikit-learn, PyTorch, and XGBoost are commonly used alongside these resources.
Scalable Machine Learning Platform
A fast, scalable, open-source machine learning and artificial intelligence platform. H2O supports widely used statistical and machine learning algorithms including gradient boosted machines, random forests, deep learning, and more with Python and R APIs.
Distributed Machine Learning
An environment for quickly creating scalable, performant machine learning applications. Mahout provides mathematically expressive Scala DSL and supports Apache Spark and Apache Flink backends for distributed linear algebra operations.
Spark's Machine Learning Library
Apache Spark's scalable machine learning library consisting of common learning algorithms and utilities including classification, regression, clustering, collaborative filtering, and dimensionality reduction. MLlib integrates seamlessly with Spark's data processing pipelines.
Posts, comments, and upvotes via PRAW. Great starting point for sentiment analysis and NLP.
Access GPT language models, embeddings, and image generation tools from OpenAI. Commonly used in dat
A simplified Wolfram Alpha endpoint that returns concise plain-text answers to factual and computati
Generate realistic synthetic user profiles including names, addresses, photos, email addresses, and
An open-source music encyclopaedia API providing structured data on artists, albums, recordings, lab
A free, open-source database API of breweries worldwide with details on beer types, locations, addre
Retrieve real-time and historical air quality measurements including PM2.5, PM10, ozone, NO2, and CO
A curated repository of 600+ datasets covering classification, regression, clustering, and time-seri
Thousands of publicly available datasets hosted on GitHub repositories covering social media, financ
Data.gov hosts 300,000+ datasets from US federal agencies covering health, education, environment, a
Data.gov.uk provides datasets from UK central and local government covering crime, transport, planni
Regular XML snapshots of all Wikipedia articles, talk pages, and revision histories available for bu
NOAA platform provides access to a vast collection of climate-related datasets, including historical
The FEC provides access to campaign finance data, including information on political contributions,
The Amazon Customer Reviews dataset on AWS Open Data contains 130+ million product reviews across 40
The US Federal Aviation Administration publishes datasets on aircraft registrations, pilot certifica
The Kaggle COVID-19 Dataset, curated by the Allen Institute for AI, aggregates a comprehensive colle
Google Cloud hosts petabyte-scale public datasets including genomics, satellite imagery, financial m
The GTD, maintained by the National Consortium for the Study of Terrorism and Responses to Terrorism
The United Nations Development Programme publishes datasets on the Human Development Index, poverty
Eurostat, the statistical office of the European Union, offers a comprehensive database of statistic
This a collaborative database of food products from around the world, containing information on ingr
The Hugging Face Datasets library provides programmatic access to 50,000+ NLP, computer vision, and
Natural Earth provides public domain map datasets at various scales, covering physical and cultural
WFP provides datasets on food security, hunger, malnutrition, food aid distribution, humanitarian as
The GUO offers datasets on urbanization, urban population growth, city demographics, slum population