Category: Deep Learning Tutorials


  • Standardization Vs Normalization

    Standardization Vs Normalization What Is Standardization ? What Is Normalization ? Why Were They Invented ? When To Use What ? How Standardization Helps In Sigmoid Satuaration In Neural Net ?

    Read More

  • General Contextual Embedding Vs Task Specific Contextual Embedding ?

    General Contextual Embedding Vs Task Specific Contextual Embedding Reason 1: Reason 2:

    Read More

  • Problem With Word2Vec Algorithm.

    Problem With Word2Vec

    Read More

  • TF-IDF All Doubts Cleared .

    TF-IDF All Doubts Cleared First tell me about term, document and corpus difference Explain Me About TF-IDF. Can the document and the corpus be totally different context, document can be from computer science and corpus can contain astrology ? While calculating TF , which document i need to check the presence of the word out of 1000 docs Application Of TF IDF where actually it is being used Why Its Called Inverse In IDF. Can TF-IDF Value Used For Sentence Embedding Purpose?

    Read More

  • What Is Gradient & How Its Calculated ?

    What Is Gradient & How Its Calculated ? What Is Positive Slope & Negative Slope ? Interms Of Neural Network? How we have established the relation between error function and the weight ?

    Read More

  • Deep Learning – What Is Early Stopping ?

    Deep Learning – What Is Early Stopping ?

    Deep Learning – What Is Early Stopping ? Table Of Contents: What Is Early Stopping ? Why Is Early Stopping Is Needed ? How Early Stopping Works ? Benefits Of Early Stopping . Visual Representation. Hyperparameter : Patience . (1) What Is Early Stopping ? (2) Why Is Early Stopping Needed ? (3) How Early Stopping Works ? (4) Benefits of Early Stopping . (5) Visual Representation . (6) Hyperparameter: Patience (7) Implementation in Keras (TensorFlow) from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense from tensorflow.keras.callbacks import EarlyStopping # 1. Build the model model = Sequential([ Dense(128, activation='relu', input_shape=(input_dim,)), Dense(64,

    Read More

  • Transformer – Prediction Process

    Transformer – Prediction Process

    Transformer – Transformer Prediction Table Of Contents: Prediction Setup Of Transformer. Step By Step Flow Of Input Sentence, “We Are Friends!”. Decoder Processing For Other Timestep (1) Prediction Setup For Transformer. Input Dataset: For simplicity we will take this 3 rows as input but in reality we will have thousands of rows as input. We will use these dataset to train our Transformer model. Query Sentence: We will pass this sentence for translation, Sentence = “We Are Friends !” (2) Step By Step Flow Of Input Sentence, “We Are Friends!”. Transformer is mainly divides into Encoder and Decoder. Encoder will

    Read More

  • Transformer – Decoder Architecture

    Transformer – Decoder Architecture

    Transformer – Decoder Architecture Table Of Contents: What Is The Work Of Decoder In Transformer ? Overall Decoder Architecture. Understanding Decoder Work Flow With An Example. Understanding Decoder 2nd Part. (1) What Is The Work Of Decoder In Transformer ? In a Transformer model, the Decoder plays a crucial role in generating output sequences from the encoded input. It is mainly used in sequence-to-sequence (Seq2Seq) tasks such as machine translation, text generation, and summarization. (2) Overall Decoder Architecture. In the original paper of Transformer we have 6 decoder module connected in series. The output from one decoder module will be

    Read More

  • Transformers – Cross Attention

    Transformers – Cross Attention

    Transformer – Cross Attention Table Of Contents: Where Is Cross Attention Block Is Applied In Transformers? What Is Cross Attention ? How Cross Attention Works? Where We Use Cross Attention Mechanism. (1) Where Is Cross Attention Block Is Applied In Transformers? In the diagram above you can see that, the Multi-Head Attention is known as “Cross Attention”. The difference to the other “Multi Head Attention” block is that for other the 3 inputs Query, Key and Value vectors are generated from a single source but in this Cross Attention block the Query vector is coming from the Decoder block and

    Read More

  • Transformer – Masked Self Attention

    Transformer – Masked Self Attention

    Transformer – Masked Self Attention Table Of Contents: Transformer Decoder Definition. What Is Autoregressive Model? Lets Prove The Transformer Decoder Definition. How To Implement The Parallel Processing Logic While Training The Transformer Decoder? Implementing Masked Self Attention. (1) Transformer Decoder Definition From this above definition we can we can understand that the Transformer behaves Autoregressive while prediction and Non Auto Regressive while training. This is displayed in the diagram below. (2) What Is Autoregressive Model? Suppose you are making a Machine Learning model which work is predict the stock price,  Monday it has predicted 29, Tuesday = 25 for to

    Read More