Glossary
The words of artificial intelligence, explained plainly.
158 termsA
- MetricAccuracyProportion of correct predictions out of total predictions made.
- ConceptActivation FunctionA mathematical function applied to a neuron's output to introduce non-linearity into the network.
- ArchitectureActivation LayerA layer that applies a non-linear activation function element-wise to its input.
- AlgorithmAdagradAdaptive learning rate optimizer that adapts rates per parameter based on historical gradients.
- AlgorithmAdam OptimizerAn adaptive learning rate optimization algorithm combining momentum and RMSprop.
- TechniqueAdversarial TrainingTraining technique improving model robustness by including adversarial examples.
- ArchitectureAlexNetDeep CNN that won ImageNet 2012, pioneering deep learning in computer vision.
- ConceptAlgorithmA step-by-step procedure or formula for solving a problem or performing a task.
- ConceptArtificial Intelligence (AI)Field of computer science focused on creating systems capable of performing tasks requiring human intelligence.
- TechniqueAttention MechanismA technique allowing models to focus on specific parts of the input when producing output.
- ConceptAttention ScoreComputed similarity between query and key vectors before softmax normalization in attention.
- ConceptAttention WeightA scalar value indicating how much focus to place on a specific part of the input when producing output.
- MetricAUC-ROCArea under ROC curve, measuring binary classifier quality across all thresholds.
- ArchitectureAutoencoderA neural network trained to reconstruct its input, learning compressed representations in the process.
- TechniqueAverage PoolingA pooling operation that computes the average value from each window of the feature map.
B
- TrainingBackpropagationAn algorithm for training neural networks by calculating gradients of the loss function with respect to weights.
- AlgorithmBatch Gradient DescentA gradient descent variant that computes gradients using the entire training dataset in each iteration.
- TechniqueBatch NormalizationNormalizes layer inputs using batch statistics to stabilize and accelerate training.
- ConceptBatch SizeNumber of training examples processed together in one forward/backward pass.
- AlgorithmBeam SearchDecoding algorithm keeping top B most likely sequences at each step.
- ConceptBias TermAn additional learnable parameter in neural networks that allows shifting the activation function.
- MetricBLEU ScoreMetric for machine translation quality comparing n-gram overlap between generated and reference translations.
C
- TechniqueCausal MaskingMasking technique preventing attention to future positions in autoregressive models.
- TechniqueChain-of-Thought PromptingPrompting technique encouraging LLMs to show intermediate reasoning steps before answering.
- TaskClassificationA supervised learning task where the goal is to predict discrete class labels for input data.
- ModelCLIP (Contrastive Language-Image Pre-training)A multimodal model trained to understand relationships between images and text.
- TaskClusteringAn unsupervised learning task that groups similar data points together based on their features.
- ArchitectureCNN (Convolutional Neural Network)A deep learning architecture specialized for processing grid-like data such as images, using convolutional layers.
- DatasetCOCO (Common Objects in Context)Large-scale dataset for object detection, segmentation, and captioning with 330k images.
- FieldComputer VisionA field of AI enabling computers to derive meaningful information from visual inputs like images and videos.
- ConceptConfusion MatrixA table used to evaluate classification model performance by showing true vs predicted classes.
- ParadigmContrastive LearningSelf-supervised learning contrasting positive pairs against negative pairs.
- ConceptConvolutionA mathematical operation that slides a filter/kernel over input data to extract features.
- ArchitectureConvolutional LayerA layer in CNNs that applies convolution operations to extract spatial features from input data.
- TechniqueCross-ValidationResampling technique evaluating model performance by splitting data into multiple train-test folds.
D
- TechniqueData AugmentationTechniques to artificially increase training data size by creating modified versions of existing data.
- ConceptDatasetA collection of data examples used for training, validation, or testing machine learning models.
- ConceptDeep LearningA subset of machine learning using neural networks with multiple layers to learn hierarchical representations of data.
- TechniqueDepthwise Separable ConvolutionFactorizes standard convolution into depthwise and pointwise convolutions for efficiency.
- TechniqueDice LossLoss function based on Dice coefficient, commonly used for segmentation tasks.
- ArchitectureDiffusion ModelsGenerative models that learn to create data by reversing a gradual noising process.
- TechniqueDilated ConvolutionConvolution with gaps between kernel elements, increasing receptive field without adding parameters.
- TechniqueDimensionality ReductionTechniques to reduce the number of features in data while preserving important information.
- TechniqueDropoutA regularization technique that randomly deactivates neurons during training to prevent overfitting.
E
- TechniqueEarly StoppingA regularization technique that stops training when validation performance stops improving.
- TechniqueEmbeddingA dense vector representation of discrete entities (words, images) in a continuous space.
- ArchitectureEncoder-DecoderAn architecture where an encoder processes input into a representation and a decoder generates output from it.
- ConceptEpochOne complete pass through the entire training dataset during training.
- ConceptErrorThe difference between predicted output and true label, indicating model's mistakes.
- ConceptExploding GradientA problem where gradients become extremely large during training, causing unstable updates and divergence.
F
- MetricF1 ScoreThe harmonic mean of precision and recall, providing a single balanced metric.
- ConceptFeatureAn individual measurable property or characteristic of data used as input to a model.
- TechniqueFeature ExtractionThe process of transforming raw data into numerical features that machine learning models can process.
- ConceptFeature MapThe output of applying a convolution filter to an input, representing detected features.
- ConceptFew-Shot LearningA model's ability to learn from a small number of examples, typically 1-10 examples per class.
- ConceptFilter / KernelA small matrix of learnable weights that slides over input during convolution to detect specific features.
- TrainingFine-tuningThe process of adapting a pre-trained model to a specific task by continuing training on task-specific data.
- TechniqueFocal LossLoss function addressing class imbalance by down-weighting easy examples.
- ConceptFoundation ModelLarge-scale models trained on broad data that can be adapted to a wide range of downstream tasks.
- ArchitectureFully Connected LayerA neural network layer where every neuron is connected to every neuron in the previous and next layers.
G
- ArchitectureGAN (Generative Adversarial Network)A framework where two neural networks compete: a generator creates fake data and a discriminator tries to distinguish real from fake.
- ModelGPT (Generative Pre-trained Transformer)A family of large language models developed by OpenAI that use transformer architecture for text generation.
- ConceptGradientDirection and magnitude of steepest increase in loss function with respect to parameters.
- AlgorithmGradient DescentAn optimization algorithm that iteratively adjusts parameters to minimize a loss function by following the gradient.
- TechniqueGroup NormalizationNormalizes by dividing channels into groups and normalizing within each group.
- ArchitectureGRU (Gated Recurrent Unit)Simplified RNN variant with gating mechanisms, similar to LSTM but fewer parameters.
H
I
- DatasetImageNetLarge-scale image dataset with 14M images across 20k categories, used for ILSVRC competition.
- ArchitectureInception ModuleCNN building block applying multiple filter sizes in parallel and concatenating results.
- ConceptInferenceThe process of using a trained model to make predictions on new data.
- ConceptInputData fed into a model or neural network for processing.
- ArchitectureInput LayerThe first layer of a neural network that receives raw input data.
- TechniqueInstance NormalizationNormalizes each sample independently, commonly used in style transfer.
L
- ConceptLabelThe target output or ground truth associated with a training example in supervised learning.
- TrainingLearning RateA hyperparameter controlling how much model weights are updated during training.
- ConceptLearning RateHyperparameter controlling how much to adjust model weights during training.
- TechniqueLearning Rate ScheduleA strategy for adjusting the learning rate during training to improve convergence and performance.
- ModelLLM (Large Language Model)A neural network with billions of parameters trained on massive text datasets to understand and generate human language.
- TechniqueLoRA (Low-Rank Adaptation)Parameter-efficient fine-tuning adding trainable low-rank matrices to frozen weights.
- ConceptLossA measure of how wrong the model's predictions are, used to guide training.
- ConceptLoss FunctionA function that measures the difference between predicted and actual values, guiding model optimization.
M
- TaskMachine TranslationAutomatic translation of text from one language to another.
- MetricMAE (Mean Absolute Error)Loss function measuring average absolute difference between predicted and actual values.
- TaskMasked Language ModelingPre-training task where random tokens are masked and model predicts them from context.
- ConceptMatrixA 2D array of numbers arranged in rows and columns.
- TechniqueMax PoolingA pooling operation that takes the maximum value from each window of the feature map.
- TechniqueMixupData augmentation creating synthetic examples by mixing pairs of training samples.
- ArchitectureMobileNetEfficient CNN architecture using depthwise separable convolutions for mobile deployment.
- ConceptModelA mathematical representation learned from data that makes predictions or decisions.
- TechniqueMomentumAn optimization technique that accelerates gradient descent by accumulating a velocity vector in directions of persistent reduction in the loss.
- MetricMSE (Mean Squared Error)Loss function measuring average squared difference between predicted and actual values.
- ConceptMultimodal AIAI systems that can process and relate information from multiple modalities like text, images, audio, and video.
N
- TaskNER (Named Entity Recognition)NLP task identifying and classifying named entities (persons, organizations, locations) in text.
- ArchitectureNeural NetworkA computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) that process information.
- ConceptNeuronBasic computational unit in neural networks that receives inputs, applies weights and activation, produces output.
- TaskNext Sentence PredictionPre-training task predicting whether sentence B follows sentence A.
- TechniqueNucleus Sampling (Top-P)Text generation sampling from smallest set of tokens whose cumulative probability exceeds threshold P.
O
- TaskObject DetectionA computer vision task that identifies and localizes objects within an image using bounding boxes.
- ConceptOptimizationThe process of adjusting model parameters to minimize the loss function and improve performance.
- ConceptOutputThe result produced by a model after processing input data.
- ArchitectureOutput LayerThe final layer of a neural network that produces predictions or outputs.
- ConceptOverfittingWhen a model learns training data too well, including noise, resulting in poor generalization to new data.
P
- TechniquePaddingAdding extra pixels around the border of input data to control output size in convolution operations.
- ConceptParameterLearnable values in a model that are optimized during training (weights and biases).
- ArchitecturePerceptronSimplest neural network with single layer, binary classifier invented in 1950s.
- MetricPerplexityMeasurement of how well a probability model predicts a sample, for evaluating language models.
- ArchitecturePooling LayerA downsampling layer in CNNs that reduces spatial dimensions while retaining important features.
- TechniquePositional EncodingA technique to inject position information into transformer inputs since transformers lack inherent sequence order.
- ConceptPrecision and RecallMetrics for classification: Precision is correct positives / predicted positives; Recall is correct positives / actual positives.
- ConceptPredictionThe output produced by a trained model when given new input data.
- TechniquePrompt EngineeringThe practice of designing effective text prompts to guide large language models toward desired outputs.
- TechniquePruningRemoving unnecessary weights/neurons to reduce model size and computational cost.
Q
- TechniqueQuantizationReducing precision of weights/activations to lower memory and computation.
- ConceptQuery, Key, Value (QKV)Three vectors used in attention mechanisms to compute weighted combinations of input elements.
- TaskQuestion AnsweringNLP task where model extracts or generates answers to questions based on context.
R
- TechniqueRAG (Retrieval-Augmented Generation)A technique that enhances LLM outputs by retrieving relevant information from external knowledge bases.
- ConceptReceptive FieldThe region of input space that affects a particular neuron's activation in a neural network.
- TaskRegressionA supervised learning task where the goal is to predict continuous numerical values.
- TechniqueRegularizationTechniques to prevent overfitting by adding constraints or penalties to the model during training.
- ParadigmReinforcement LearningA machine learning paradigm where agents learn by interacting with an environment and receiving rewards or penalties.
- ConceptReLU (Rectified Linear Unit)An activation function that outputs the input if positive, otherwise zero: f(x) = max(0, x).
- ArchitectureResNet (Residual Network)A CNN architecture that uses residual connections (skip connections) to enable training of very deep networks.
- AlgorithmRMSpropAdaptive learning rate optimization algorithm using moving average of squared gradients.
- ArchitectureRNN (Recurrent Neural Network)A neural network architecture designed for sequential data, with connections that loop back to previous states.
S
- ConceptScalarA single numerical value, a zero-dimensional tensor.
- TechniqueSelf-AttentionAn attention mechanism used in deep learning models that allows a neural network to weigh the importance of different parts of an input relative to each other.
- ParadigmSelf-Supervised LearningLearning paradigm where models create supervision signal from unlabeled data.
- TaskSemantic SegmentationA computer vision task that assigns a class label to every pixel in an image.
- TaskSentiment AnalysisNLP task determining emotional tone or opinion expressed in text.
- AlgorithmSGD (Stochastic Gradient Descent)A gradient descent variant that updates weights using gradients from a single random training example at a time.
- ConceptSigmoid FunctionAn activation function that maps inputs to values between 0 and 1: f(x) = 1/(1 + e^(-x)).
- ConceptSoftmax FunctionAn activation function that converts a vector of values into a probability distribution summing to 1.
- DatasetSQuAD (Stanford Question Answering Dataset)Reading comprehension dataset with 100k+ questions on Wikipedia articles.
- ConceptStrideThe number of pixels by which a filter moves across the input during convolution or pooling operations.
- ParadigmSupervised LearningA machine learning paradigm where models learn from labeled training data with input-output pairs.
T
- TechniqueTemperature (Sampling)Parameter controlling randomness in text generation by scaling logits before softmax.
- ConceptTest DataData held out for final evaluation of a trained model, never seen during training.
- TaskText SummarizationNLP task condensing long text while preserving key information.
- TechniqueTokenizationThe process of breaking text into smaller units (tokens) like words, subwords, or characters for processing.
- TechniqueTop-K SamplingText generation technique sampling from only the K most likely next tokens.
- ConceptTraining DataThe subset of data used to train a machine learning model.
- TrainingTransfer LearningA technique where knowledge learned from one task is applied to a different but related task, reducing training time and data requirements.
- ArchitectureTransformerA neural network architecture based entirely on attention mechanisms, without recurrent or convolutional layers.
- TechniqueTriplet LossLoss function learning embeddings by minimizing distance between anchor-positive and maximizing anchor-negative.
U
V
- ArchitectureVAE (Variational Autoencoder)A generative model that learns a probabilistic latent space representation of data.
- ConceptValidation SetA portion of data held out during training to tune hyperparameters and prevent overfitting.
- ConceptVanishing GradientA problem in deep networks where gradients become extremely small, preventing effective learning in early layers.
- ConceptVectorAn ordered array of numbers representing a point in multi-dimensional space.
- ArchitectureVGGDeep CNN architecture using small 3x3 filters throughout, emphasizing depth.
- ArchitectureViT (Vision Transformer)Transformer architecture adapted for computer vision by treating image patches as tokens.
W
- ConceptWeightLearnable parameters in neural networks that determine the strength of connections between neurons.
- TechniqueWeight InitializationMethods for setting initial values of neural network weights before training begins.
- TechniqueWord EmbeddingDense vector representations of words that capture semantic and syntactic relationships.
Y
Z
No terms found.