Viva prep · Real questions · Student experiences Enroll in Bootcamp

Search viva questions

Search by question text across the whole library — then narrow with subject, level, or proctor.

Question sets

  1. 1
    Level 2 Viva Advice
    Times asked 1
  2. 2
    Be extremely clear with your notebook.
    Times asked 1
  3. 3
    Be able to explain every command, parameter, and implementation you have used.
    Times asked 1
  4. 4
    Understand the logic, intuition, and mathematical concepts behind every step in your notebook.
    Times asked 1
  1. 1
    Explain your notebook/project.
    Times asked 21
  2. 2
    How does Logistic Regression work?
    Times asked 3
  3. 3
    What is One-Hot Encoding? Give an example.
    Times asked 1
  4. 4
    Did you use Hyperparameter Tuning?
    Times asked 1
  5. 5
    Did you use Pipeline?
    Times asked 1
  1. 1
    Show your ID card.
    Times asked 23
  2. 2
    Explain your notebook/project.
    Times asked 21
  3. 3
    Explain your notebook as a story.
    Times asked 2
  4. 4
    Explain your EDA.
    Times asked 8
  5. 5
    Explain the graphs you created.
    Times asked 1
  6. 6
    What insights did you draw from the graphs?
    Times asked 1
  7. 7
    Difference between a bar plot and a histogram.
    Times asked 1
  8. 8
    What does a pair plot show?
    Times asked 1
  9. 9
    What is the shaded region in a regression plot?
    Times asked 1
  10. 10
    Why is the correlation between two features zero?
    Times asked 1
  11. 11
    How does correlation affect model predictions?
    Times asked 1
  12. 12
    What does errors='coerce' do while converting dates?
    Times asked 1
  13. 13
    Explain your preprocessing steps.
    Times asked 2
  14. 14
    Explain your train-test split.
    Times asked 1
  15. 15
    What happens if the train-test split ratio changes?
    Times asked 1
  16. 16
    Explain your encoding.
    Times asked 2
  17. 17
    Explain your scaling.
    Times asked 2
  18. 18
    Explain your feature engineering.
    Times asked 6
  19. 19
    Why did you use ColumnTransformer?
    Times asked 1
  20. 20
    Why didn't you use Pipeline?
    Times asked 1
  21. 21
    Have you used user-defined functions?
    Times asked 1
  22. 22
    Explain your hyperparameter tuning.
    Times asked 3
  23. 23
    Why did you use RandomizedSearchCV?
    Times asked 1
  24. 24
    Difference between GridSearchCV and RandomizedSearchCV.
    Times asked 5
  25. 25
    How many times will RandomizedSearchCV run?
    Times asked 1
  26. 26
    What scoring metrics can be used in GridSearchCV?
    Times asked 1
  27. 27
    How do different scoring metrics affect model selection?
    Times asked 1
  28. 28
    What models did you use?
    Times asked 3
  29. 29
    Why did you choose those models?
    Times asked 3
  30. 30
    What is your baseline model?
    Times asked 1
  31. 31
    Why did you use only boosting models?
    Times asked 1
  32. 32
    Difference between Random Forest and XGBoost.
    Times asked 2
  33. 33
    What is Random Forest?
    Times asked 1
  34. 34
    Why didn't you choose Random Forest?
    Times asked 1
  35. 35
    What is XGBoost?
    Times asked 2
  36. 36
    What is LightGBM?
    Times asked 1
  37. 37
    What are the important LightGBM parameters?
    Times asked 1
  38. 38
    What is early_stopping?
    Times asked 1
  39. 39
    After early stopping, when is the best score reported?
    Times asked 1
  40. 40
    What are min_samples_split and min_samples_leaf?
    Times asked 1
  41. 41
    If a node has 7 samples, will it split?
    Times asked 1
  42. 42
    Which of your models are parametric?
    Times asked 1
  43. 43
    Which of your models are non-parametric?
    Times asked 1
  44. 44
    Why is Logistic Regression called "Regression" even though it performs classification?
    Times asked 1
  45. 45
    What is MLPClassifier?
    Times asked 1
  46. 46
    What is Ridge Regression?
    Times asked 1
  47. 47
    What is Lasso Regression?
    Times asked 1
  48. 48
    Which regression technique can be used for feature selection?
    Times asked 1
  49. 49
    What is Stacking?
    Times asked 1
  50. 50
    How does stacking work?
    Times asked 1
  51. 51
    How does a boosting model learn?
    Times asked 1
  52. 52
    How does a Ridge meta-model combine XGBoost, CatBoost, and LightGBM?
    Times asked 1
  53. 53
    How did you choose ensemble weights?
    Times asked 1
  54. 54
    Why did your ensemble use those weights?
    Times asked 1
  55. 55
    What is weak learner predictive power?
    Times asked 1
  56. 56
    How much data is used in Bagging?
    Times asked 1
  57. 57
    What is F1-Score?
    Times asked 3
  58. 58
    Write the formula for F1-Score.
    Times asked 2
  59. 59
    Does F1-Score give equal importance to Precision and Recall?
    Times asked 1
  60. 60
    Why did you use F1 Macro?
    Times asked 1
  61. 61
    Why didn't you use Accuracy?
    Times asked 1
  62. 62
    Explain the result comparison graph.
    Times asked 1
  63. 63
    Explain your model performance comparison.
    Times asked 1
  64. 64
    Coding: Load the Iris dataset and print the feature matrix.
    Times asked 1
  65. 65
    Coding: Load the Iris dataset and print the target values.
    Times asked 1
  66. 66
    Coding: Load the Iris dataset and separate features and target.
    Times asked 1
  67. 67
    Coding: Load the Diabetes dataset and print the feature matrix.
    Times asked 1
  68. 68
    Coding: Load the Breast Cancer dataset and print the target names.
    Times asked 1
  69. 69
    Coding: Load the California Housing dataset and train a HistGradientBoostingRegressor.
    Times asked 1
  70. 70
    Coding: Draw a histogram using the Diabetes dataset.
    Times asked 1
  71. 71
    Coding: Plot a bar graph comparing models and their scores.
    Times asked 1
  72. 72
    Coding: Plot the target variable after applying log transformation.
    Times asked 1
  73. 73
    Coding: Implement GridSearchCV or RandomizedSearchCV.
    Times asked 1
  74. 74
    Coding: Import a dummy/sample dataset.
    Times asked 1
  75. 75
    Coding: Implement Lasso Regression.
    Times asked 1
  76. 76
    Coding: Print the first few rows of the Iris dataset.
    Times asked 1
  77. 77
    Coding: Use fit_transform() output as a Pandas DataFrame using set_output(transform="pandas").
    Times asked 1
  1. 1
    Introduction.
    Times asked 9
  2. 2
    Explain your notebook/project.
    Times asked 21
  3. 3
    Show and explain each section of your notebook.
    Times asked 1
  4. 4
    No additional theory questions were asked.
    Times asked 2
  1. 1
    Explain your notebook/code.
    Times asked 1
  2. 2
    Questions based on the selector used.
    Times asked 1
  3. 3
    Questions based on the scaler used.
    Times asked 1
  1. 1
    Introduction.
    Times asked 9
  2. 2
    Show your ID card.
    Times asked 23
  3. 3
    Explain your notebook/project.
    Times asked 21
  4. 4
    Explain your notebook as a story.
    Times asked 2
  5. 5
    Explain your project flow from start to finish.
    Times asked 1
  6. 6
    Explain the machine learning models you used.
    Times asked 1
  7. 7
    Explain the preprocessing pipeline you used.
    Times asked 1
  8. 8
    Explain the hyperparameter tuning process.
    Times asked 1
  9. 9
    What is the cv parameter in GridSearchCV/RandomizedSearchCV?
    Times asked 1
  10. 10
    How does Cross Validation work?
    Times asked 2
  11. 11
    What is K-Fold Cross Validation?
    Times asked 2
  12. 12
    Why is a validation set needed?
    Times asked 1
  13. 13
    How many iterations/model fits will be performed based on your hyperparameter tuning?
    Times asked 1
  14. 14
    Difference between GridSearchCV and RandomizedSearchCV.
    Times asked 5
  15. 15
    How do you identify whether a model is overfitting or underfitting?
    Times asked 1
  16. 16
    How do you reduce overfitting?
    Times asked 2
  17. 17
    What happens if the correlation between features is high?
    Times asked 1
  18. 18
    How many models did you try?
    Times asked 2
  19. 19
    Did you try different preprocessing techniques?
    Times asked 1
  20. 20
    Did you try different models?
    Times asked 1
  21. 21
    What resources did you use?
    Times asked 2
  22. 22
    What other approaches did you try?
    Times asked 1
  23. 23
    What difficulties did you face while doing the project?
    Times asked 1
  24. 24
    Have you used any sklearn Pipelines?
    Times asked 1
  25. 25
    Explain your missing value handling.
    Times asked 1
  26. 26
    How did you handle numerical missing values?
    Times asked 2
  27. 27
    How did you handle class imbalance?
    Times asked 3
  28. 28
    Why did you use One-Hot Encoding?
    Times asked 1
  29. 29
    Have you used scaling?
    Times asked 1
  30. 30
    Why did you use StandardScaler?
    Times asked 2
  31. 31
    How does StandardScaler work?
    Times asked 2
  32. 32
    Difference between StandardScaler and other scaling techniques.
    Times asked 1
  33. 33
    What is scaling?
    Times asked 5
  34. 34
    Why didn't you use scaling?
    Times asked 2
  35. 35
    What is the learning rate parameter?
    Times asked 1
  36. 36
    Which scoring metric did you use to compare models?
    Times asked 1
  37. 37
    Why did you use F1-Score, Precision, and Recall instead of RMSE?
    Times asked 1
  38. 38
    What happens if RMSE is used for Logistic Regression?
    Times asked 1
  39. 39
    Explain the ROC Curve.
    Times asked 1
  40. 40
    Explain the Confusion Matrix.
    Times asked 1
  41. 41
    Point out True Positive (TP) and True Negative (TN) in the Confusion Matrix.
    Times asked 1
  42. 42
    What is log1p?
    Times asked 1
  43. 43
    What is the mean and standard deviation after StandardScaler?
    Times asked 1
  44. 44
    Coding: Implement RandomizedSearchCV for Logistic Regression.
    Times asked 1
  45. 45
    Coding: Train an ElasticNet model and evaluate its performance.
    Times asked 1
  46. 46
    Coding: Train a Linear Regression model on the raw (non-preprocessed) dataset.
    Times asked 1
  47. 47
    Coding: Load the training dataset and print the first 6/7/10 rows.
    Times asked 1
  48. 48
    Coding: Filter rows where ManufactureYear is greater than a given year.
    Times asked 1
  49. 49
    Coding: Print unique/value counts of VendorPartnerID.
    Times asked 1
  50. 50
    Coding: Find the proportion of each category (cardinality) in a feature.
    Times asked 1
  51. 51
    Coding: Load the Iris dataset and train a Ridge Regression model.
    Times asked 1
  52. 52
    Coding: Load the California Housing dataset.
    Times asked 2
  1. 1
    Tell me about yourself.
    Times asked 6
  2. 2
    What is your Kaggle leaderboard score/rank?
    Times asked 1
  3. 3
    Explain your notebook/project.
    Times asked 21
  4. 4
    Explain your EDA.
    Times asked 8
  5. 5
    How many graphs did you create?
    Times asked 1
  6. 6
    Explain each graph you created.
    Times asked 1
  7. 7
    What observations did you make from your graphs?
    Times asked 1
  8. 8
    How do you hide/show values in a correlation heatmap?
    Times asked 1
  9. 9
    Explain your boxplot.
    Times asked 1
  10. 10
    Explain your density regression plot.
    Times asked 1
  11. 11
    What feature engineering did you perform?
    Times asked 2
  12. 12
    Why did you perform that feature engineering?
    Times asked 1
  13. 13
    How did you handle missing values?
    Times asked 4
  14. 14
    How did you handle numerical missing values?
    Times asked 2
  15. 15
    How did you handle categorical missing values?
    Times asked 1
  16. 16
    How did you handle disguised missing values (None, NaN, Unknown, Unspecified)?
    Times asked 1
  17. 17
    Why did you preserve the "Unknown" category?
    Times asked 1
  18. 18
    What numerical transformers did you use?
    Times asked 1
  19. 19
    What categorical transformers did you use?
    Times asked 1
  20. 20
    Why did you use those transformers?
    Times asked 1
  21. 21
    What is Data Leakage?
    Times asked 1
  22. 22
    When does Data Leakage occur?
    Times asked 1
  23. 23
    How did you prevent Data Leakage?
    Times asked 1
  24. 24
    What is an N-gram?
    Times asked 1
  25. 25
    What is Unigram?
    Times asked 1
  26. 26
    What is Bigram?
    Times asked 1
  27. 27
    What is a 6-gram, 8-gram, or 11-gram?
    Times asked 1
  28. 28
    What is TF-IDF?
    Times asked 2
  29. 29
    Difference between word-level and character-level TF-IDF.
    Times asked 1
  30. 30
    What does TF-IDF Vectorizer do?
    Times asked 1
  31. 31
    Explain the parameters of TF-IDF Vectorizer.
    Times asked 1
  32. 32
    Why did you choose 50,000 TF-IDF features?
    Times asked 1
  33. 33
    What models did you use?
    Times asked 3
  34. 34
    Why did you choose those models?
    Times asked 3
  35. 35
    What is your best model?
    Times asked 1
  36. 36
    Why is it your best model?
    Times asked 2
  37. 37
    What is an Ensemble model?
    Times asked 1
  38. 38
    Explain your final modeling approach.
    Times asked 1
  39. 39
    What hyperparameter tuning did you perform?
    Times asked 1
  40. 40
    Which hyperparameters did you tune?
    Times asked 2
  41. 41
    How many model fits will GridSearchCV perform?
    Times asked 1
  42. 42
    How many model fits will RandomizedSearchCV perform?
    Times asked 1
  43. 43
    Difference between GridSearchCV and RandomizedSearchCV.
    Times asked 5
  44. 44
    How do you know how many fits will be performed?
    Times asked 1
  45. 45
    What is n_estimators?
    Times asked 2
  46. 46
    Why did you choose n_estimators = 100?
    Times asked 1
  47. 47
    What happens if n_estimators = 2000?
    Times asked 1
  48. 48
    What is max_depth?
    Times asked 1
  49. 49
    What is learning_rate?
    Times asked 1
  50. 50
    What is n_iter?
    Times asked 1
  51. 51
    What would you do if the model overfits because of max_depth?
    Times asked 1
  52. 52
    How do you prevent overfitting?
    Times asked 1
  53. 53
    How do you improve an underfitting model?
    Times asked 1
  54. 54
    What is IQR?
    Times asked 1
  55. 55
    How does outlier detection help?
    Times asked 1
  56. 56
    How does a boxplot help detect outliers?
    Times asked 1
  57. 57
    How does ColumnTransformer work?
    Times asked 1
  58. 58
    Why didn't you use scaling?
    Times asked 2
  59. 59
    Should SGD be used without scaling?
    Times asked 1
  60. 60
    What is SGD?
    Times asked 1
  61. 61
    Why did you use SGD?
    Times asked 1
  62. 62
    Difference between Accuracy, Precision, Recall, and F1-Score.
    Times asked 1
  63. 63
    Why did you use F1-Score?
    Times asked 1
  64. 64
    What is RMSLE?
    Times asked 1
  65. 65
    What is Blended Score?
    Times asked 1
  66. 66
    Why did your RMSLE score decrease after blending?
    Times asked 1
  67. 67
    What is a Decision Tree?
    Times asked 1
  68. 68
    What algorithm do XGBoost, LightGBM, and CatBoost use?
    Times asked 1
  69. 69
    Difference between XGBoost, LightGBM, and CatBoost.
    Times asked 1
  70. 70
    Are these models linear or non-linear?
    Times asked 1
  71. 71
    Explain your complete preprocessing and modeling pipeline.
    Times asked 1
  72. 72
    How does your complete pipeline work?
    Times asked 1
  73. 73
    What are custom classes?
    Times asked 1
  74. 74
    Why did you create custom classes?
    Times asked 1
  75. 75
    Did you use any LLMs while building the project?
    Times asked 1
  76. 76
    Coding: Impute the most frequent categorical value for each column.
    Times asked 1
  77. 77
    Coding: Perform PCA on numerical features and obtain the top two components.
    Times asked 1
  78. 78
    Coding: Perform K-Means clustering on a dataset.
    Times asked 1
  79. 79
    Coding: Load a dataset, preprocess it, perform train-test split, and build a Linear Regression model.
    Times asked 1
  80. 80
    Coding: Implement RandomizedSearchCV.
    Times asked 1
  81. 81
    Coding: Create a complete Pipeline with ColumnTransformer, Imputers, Transformers, and a model.
    Times asked 1
  1. 1
    Why do you want to do this course?
    Times asked 1
  2. 2
    What are you doing to improve your skills?
    Times asked 1
  3. 3
    Explain your notebook/project.
    Times asked 21
  4. 4
    Questions based on your project implementation.
    Times asked 1
  1. 1
    Why did you plot ROC Curve if you are already plotting f1-scores.
    Times asked 1
  1. 1
    What is the use of a correlation matrix?
    Times asked 1
  2. 2
    How do you handle imbalanced datasets?
    Times asked 2
  3. 3
    What techniques are used for handling null values?
    Times asked 1
  4. 4
    What is K-Fold Cross Validation?
    Times asked 2
  5. 5
    Which Cross Validation technique did you use?
    Times asked 1
  6. 6
    Did you drop any columns/features? Why?
    Times asked 1
  7. 7
    What is Precision?
    Times asked 2
  8. 8
    What is Recall?
    Times asked 2
  9. 9
    What is Support in a Classification Report?
    Times asked 2
  10. 10
    What is Macro Average in a Classification Report?
    Times asked 1
  11. 11
    What is Weighted Average in a Classification Report?
    Times asked 1
  12. 12
    What is the degree parameter in SVM?
    Times asked 1
  13. 13
    What is the C parameter in SVM?
    Times asked 1
  14. 14
    How does an SVM classifier work?
    Times asked 1
  15. 15
    What is an activation function?
    Times asked 1
  16. 16
    Why are activation functions used in Neural Networks?
    Times asked 1
  17. 17
    Name different activation functions.
    Times asked 2
  18. 18
    Explain the Tanh activation function.
    Times asked 1
  19. 19
    What is zero-centering?
    Times asked 1
  20. 20
    What are Neural Networks?
    Times asked 1
  21. 21
    Why did you try multiple models?
    Times asked 2
  22. 22
    What is Information Gain?
    Times asked 3
  23. 23
    How is Information Gain related to Entropy?
    Times asked 1
  24. 24
    Explain Entropy.
    Times asked 1
  25. 25
    How does a Decision Tree work?
    Times asked 2
  26. 26
    Explain the Decision Tree building process.
    Times asked 2
  27. 27
    Why did you choose a particular learning rate for LightGBM?
    Times asked 1
  28. 28
    Why did your score change after private evaluation?
    Times asked 1
  29. 29
    What EDA did you perform?
    Times asked 1
  30. 30
    What are the types of Machine Learning?
    Times asked 2
  31. 31
    Difference between Supervised and Unsupervised Learning.
    Times asked 2
  32. 32
    Difference between K-Means and K-Means++.
    Times asked 2
  33. 33
    How does clustering work?
    Times asked 1
  34. 34
    How would you use clustering on purchase/customer data?
    Times asked 1
  35. 35
    Are more features always better for training?
    Times asked 1
  36. 36
    What is the Curse of Dimensionality?
    Times asked 2
  37. 37
    How does PCA work?
    Times asked 3
  38. 38
    What is Gradient Descent?
    Times asked 2
  39. 39
    What is the Sigmoid function?
    Times asked 1
  40. 40
    Write the Sigmoid function formula.
    Times asked 2
  41. 41
    What threshold is used in Logistic Regression?
    Times asked 1
  42. 42
    Explain Logistic Regression.
    Times asked 5
  43. 43
    Can Logistic Regression handle outliers?
    Times asked 1
  44. 44
    What are the assumptions/conditions for Logistic Regression?
    Times asked 1
  45. 45
    What is Log Loss?
    Times asked 1
  46. 46
    What is the loss function of Linear Regression?
    Times asked 2
  47. 47
    Difference between RMSE and RMSLE.
    Times asked 1
  48. 48
    Why did you use RMSLE instead of RMSE?
    Times asked 1
  49. 49
    What is Naive Bayes?
    Times asked 1
  50. 50
    What is Machine Learning?
    Times asked 2
  51. 51
    Give examples of different Machine Learning algorithms.
    Times asked 1
  52. 52
    What models can be used for Sentiment Analysis?
    Times asked 1
  53. 53
    What is R² Score?
    Times asked 3
  54. 54
    How do you calculate mean and variance?
    Times asked 1
  55. 55
    How can you make a feature follow a normal distribution?
    Times asked 1
  56. 56
    Explain the T-Test.
    Times asked 1
  57. 57
    How do you compare the results obtained using RFE?
    Times asked 1
  58. 58
    Explain Bagging.
    Times asked 6
  59. 59
    Explain Boosting.
    Times asked 6
  60. 60
    What is the Bias-Variance Tradeoff?
    Times asked 3
  61. 61
    How does a Confusion Matrix work?
    Times asked 1
  62. 62
    How do you interpret a Confusion Matrix?
    Times asked 1
  63. 63
    What are feature selection techniques?
    Times asked 2
  64. 64
    Have you applied feature selection?
    Times asked 1
  65. 65
    How do you evaluate the performance of a feature selection technique?
    Times asked 1
  66. 66
    What is undersampling?
    Times asked 1
  67. 67
    What is oversampling?
    Times asked 1
  68. 68
    How do you identify whether a dataset is imbalanced?
    Times asked 1
  69. 69
    How do you choose the best model when multiple models have similar accuracy?
    Times asked 1
  70. 70
    How do you choose a model if some models take a very long time to train?
    Times asked 1
  71. 71
    How do you choose the learning rate (η) in Gradient Descent?
    Times asked 1
  72. 72
    What happens if the learning rate is too high or too low?
    Times asked 1
  73. 73
    Should the learning rate change during training?
    Times asked 1
  74. 74
    How are Precision, Recall, and F1-Score calculated?
    Times asked 1
  75. 75
    What are evaluation metrics for Classification?
    Times asked 1
  76. 76
    What are evaluation metrics for Regression?
    Times asked 1
  77. 77
    Coding: Build an SVM Pipeline (StandardScaler + SVM) on the MNIST/Digits dataset and generate the Classification Report.
    Times asked 1
  78. 78
    Coding: Compare SVM performance before and after applying PCA.
    Times asked 1
  79. 79
    Coding: Apply TF-IDF manually and calculate TF and IDF values for given documents.
    Times asked 1
  80. 80
    Coding: Calculate the mean and variance of given data points manually.
    Times asked 1
Maintaine By Lazy IITians Team + IITM BS students Regualr for More data you have