Viva prep · Real questions · Student experiences Enroll in Bootcamp

Search viva questions

Search by question text across the whole library — then narrow with subject, level, or proctor.

Question sets

  1. 1
    Show your ID card.
    Times asked 23
  2. 2
    Introduction.
    Times asked 9
  3. 3
    Explain your notebook/project.
    Times asked 21
  4. 4
    How many models did you train?
    Times asked 1
  5. 5
    On how many models did you perform Hyperparameter Tuning?
    Times asked 1
  6. 6
    Why did you perform Hyperparameter Tuning only on those models?
    Times asked 1
  7. 7
    Which model took the longest to train? Why?
    Times asked 1
  8. 8
    Explain Gradient Descent.
    Times asked 1
  9. 9
    Explain Gradient Descent on a whiteboard.
    Times asked 1
  10. 10
    Why did you use Accuracy instead of F1-Score or ROC-AUC?
    Times asked 1
  11. 11
    Draw and explain the Confusion Matrix.
    Times asked 1
  12. 12
    Write and explain the formulas for Precision, Recall, F1-Score, and Accuracy.
    Times asked 1
  13. 13
    How did you perform feature engineering on date/time features?
    Times asked 1
  14. 14
    Which hyperparameters did you tune?
    Times asked 2
  15. 15
    What time series model would you use?
    Times asked 1
  16. 16
    What more could you have done to improve the project?
    Times asked 1
  17. 17
    How did you handle class imbalance?
    Times asked 3
  18. 18
    What does the C parameter in Logistic Regression represent?
    Times asked 2
  19. 19
    Why did you use Median instead of Mean (or vice versa) for imputation?
    Times asked 1
  20. 20
    Which imputation technique would you use for a right-skewed feature?
    Times asked 1
  21. 21
    Which model did you use for the final submission?
    Times asked 1
  22. 22
    Show your Kaggle leaderboard score and rank.
    Times asked 2
  23. 23
    Show your Kaggle submissions.
    Times asked 1
  24. 24
    Questions based on your notebook implementation.
    Times asked 2
  25. 25
    Suggestions on improving the notebook/project.
    Times asked 1
  26. 26
    Asked if you have any questions for the proctor.
    Times asked 1
  1. 1
    Introduction.
    Times asked 9
  2. 2
    Explain your notebook/project.
    Times asked 21
  3. 3
    How did you handle missing (NaN) values?
    Times asked 1
  4. 4
    Which imputer should be used when a feature contains outliers?
    Times asked 1
  5. 5
    How does Logistic Regression work?
    Times asked 3
  6. 6
    What is the loss function of Logistic Regression?
    Times asked 3
  7. 7
    How does Linear Regression work?
    Times asked 1
  8. 8
    What is the loss function of Linear Regression?
    Times asked 2
  9. 9
    How does Random Forest work?
    Times asked 1
  10. 10
    How does XGBoost work?
    Times asked 2
  11. 11
    Difference between Bagging and Boosting.
    Times asked 4
  12. 12
    How can you handle an imbalanced dataset?
    Times asked 1
  13. 13
    What is the Pearson Correlation formula?
    Times asked 1
  14. 14
    Do all trees in a Random Forest receive all features?
    Times asked 1
  15. 15
    What are the types of feature reduction techniques?
    Times asked 1
  16. 16
    How does PCA work?
    Times asked 3
  17. 17
    Is PCA linear or non-linear?
    Times asked 1
  18. 18
    What are Eigenvectors?
    Times asked 1
  19. 19
    How do you calculate PCA manually?
    Times asked 1
  20. 20
    How do you calculate Eigenvectors manually?
    Times asked 1
  21. 21
    Why does L1 Regularization eliminate features?
    Times asked 1
  22. 22
    Why does L2 Regularization help prevent overfitting?
    Times asked 1
  23. 23
    In a poisonous apple detection problem, would you prioritize Precision or Recall? Why?
    Times asked 1
  1. 1
    Explain your EDA.
    Times asked 8
  2. 2
    Explain your EDA code.
    Times asked 1
  3. 3
    Suggestions to improve the notebook/code.
    Times asked 1
  4. 4
    No additional theory or coding questions were asked.
    Times asked 1
  1. 1
    Show your ID card.
    Times asked 23
  2. 2
    Introduction.
    Times asked 9
  3. 3
    Tell me about yourself.
    Times asked 6
  4. 4
    Tell me about yourself.
    Times asked 6
  5. 5
    Explain the problem statement.
    Times asked 6
  6. 6
    Explain your notebook/project.
    Times asked 21
  7. 7
    Explain your preprocessing.
    Times asked 1
  8. 8
    Explain your EDA.
    Times asked 8
  9. 9
    Explain your feature engineering.
    Times asked 6
  10. 10
    Explain your hyperparameter tuning.
    Times asked 3
  11. 11
    Explain your model comparison.
    Times asked 2
  12. 12
    What insights did you gain from your analysis?
    Times asked 1
  13. 13
    What difficulties did you face during the project?
    Times asked 1
  14. 14
    How did you overcome those difficulties?
    Times asked 1
  15. 15
    What did you learn from the project?
    Times asked 1
  16. 16
    Show an overview of your Kaggle submissions.
    Times asked 1
  17. 17
    What models did you use?
    Times asked 3
  18. 18
    Why did you use those models?
    Times asked 1
  19. 19
    Why did you use only tree/leaf-based models?
    Times asked 1
  20. 20
    Why do we perform hyperparameter tuning?
    Times asked 1
  21. 21
    What are hyperparameters?
    Times asked 1
  22. 22
    Difference between parameters and hyperparameters.
    Times asked 2
  23. 23
    How is hyperparameter tuning done?
    Times asked 1
  24. 24
    Ways to perform hyperparameter tuning.
    Times asked 1
  25. 25
    Difference between Pipeline and ColumnTransformer.
    Times asked 1
  26. 26
    Difference between Pipeline and Transformer.
    Times asked 1
  27. 27
    Why did you use StandardScaler?
    Times asked 2
  28. 28
    Which scaling method is better for handling outliers?
    Times asked 1
  29. 29
    Why do machine learning models use numerical features?
    Times asked 1
  30. 30
    What is PCA?
    Times asked 3
  31. 31
    Difference between XGBoost and LightGBM.
    Times asked 5
  32. 32
    What is overfitting?
    Times asked 5
  33. 33
    How do you reduce overfitting?
    Times asked 2
  34. 34
    What is the Bias-Variance Tradeoff?
    Times asked 3
  35. 35
    Difference between Bagging and Boosting.
    Times asked 4
  36. 36
    Which sklearn API contains OneHotEncoder?
    Times asked 1
  37. 37
    Which sklearn API contains StandardScaler?
    Times asked 1
  38. 38
    Coding: Implement the preprocessing Pipeline used in your notebook.
    Times asked 1
  39. 39
    Coding: Write a Decision Tree classifier snippet.
    Times asked 1
  40. 40
    Coding: Load the California Housing dataset.
    Times asked 2
  41. 41
    Coding: Perform a train-test split on the California Housing dataset.
    Times asked 1
  42. 42
    Coding: Load the Iris dataset.
    Times asked 1
  43. 43
    Coding: Perform a train-test split on the Iris dataset.
    Times asked 1
  44. 44
    Coding: Train a Decision Tree classifier on the Iris dataset.
    Times asked 1
  45. 45
    Coding: Print the Accuracy Score.
    Times asked 1
  46. 46
    Coding: Use the describe() function on the California Housing dataset.
    Times asked 1
  1. 1
    Show your ID card.
    Times asked 23
  2. 2
    Explain your notebook/project.
    Times asked 21
  3. 3
    Explain your EDA.
    Times asked 8
  4. 4
    Explain your univariate analysis.
    Times asked 1
  5. 5
    Explain your bivariate analysis.
    Times asked 1
  6. 6
    Explain the graphs you created and the insights from them.
    Times asked 1
  7. 7
    How did you handle missing values?
    Times asked 4
  8. 8
    What is SimpleImputer?
    Times asked 2
  9. 9
    How does SimpleImputer work?
    Times asked 1
  10. 10
    Explain your preprocessing pipeline.
    Times asked 1
  11. 11
    Have you used Pipeline?
    Times asked 2
  12. 12
    Why did you use (or not use) Pipeline?
    Times asked 1
  13. 13
    Explain your feature engineering.
    Times asked 6
  14. 14
    Did you create any new features?
    Times asked 1
  15. 15
    What encoding technique did you use?
    Times asked 1
  16. 16
    Why did you choose that encoder?
    Times asked 1
  17. 17
    Replace OrdinalEncoder with OneHotEncoder.
    Times asked 1
  18. 18
    Difference between Label Encoding and OneHot Encoding.
    Times asked 1
  19. 19
    What scaler did you use?
    Times asked 1
  20. 20
    Why did you choose MinMaxScaler over StandardScaler?
    Times asked 1
  21. 21
    Explain StandardScaler.
    Times asked 1
  22. 22
    Difference between StandardScaler and MinMaxScaler.
    Times asked 2
  23. 23
    What models did you try?
    Times asked 1
  24. 24
    Which model performed the best?
    Times asked 1
  25. 25
    Why didn't you explore linear models?
    Times asked 1
  26. 26
    Explain your model comparison.
    Times asked 2
  27. 27
    Explain the hyperparameters you tuned.
    Times asked 2
  28. 28
    What is the learning rate?
    Times asked 3
  29. 29
    Is a higher or lower learning rate better? Why?
    Times asked 1
  30. 30
    What is the Bias-Variance Tradeoff?
    Times asked 3
  31. 31
    What is Boosting?
    Times asked 2
  32. 32
    What is Gradient Boosting?
    Times asked 1
  33. 33
    What algorithm does Boosting use?
    Times asked 1
  34. 34
    What is Bagging?
    Times asked 2
  35. 35
    Difference between Bagging and Boosting.
    Times asked 4
  36. 36
    Difference between XGBoost and LightGBM.
    Times asked 5
  37. 37
    What is R² Score?
    Times asked 3
  38. 38
    Write the formula for R² Score.
    Times asked 1
  39. 39
    From which library do you import r2_score?
    Times asked 1
  40. 40
    From which library do you import OneHotEncoder?
    Times asked 1
  41. 41
    From which library do you import XGBoost?
    Times asked 1
  42. 42
    Why is Logistic Regression called "Regression" in a classification task?
    Times asked 1
  43. 43
    What is F1-Score?
    Times asked 3
  44. 44
    Write the formulas for Accuracy, Precision, and F1-Score.
    Times asked 1
  45. 45
    What are Ensemble methods?
    Times asked 1
  46. 46
    What resources did you use while building the project?
    Times asked 1
  47. 47
    Coding: Create a dataframe containing only categorical (object) columns.
    Times asked 1
  48. 48
    Coding: Apply OneHotEncoder to categorical columns with ≤10 unique values and find the new number of columns.
    Times asked 1
  49. 49
    Coding: Split the training data into numerical and categorical dataframes.
    Times asked 1
  50. 50
    Coding: Build separate preprocessing pipelines for numerical and categorical features using ColumnTransformer.
    Times asked 1
  51. 51
    Coding: Perform a train-test split.
    Times asked 1
  52. 52
    Coding: Print all numerical and non-numerical columns.
    Times asked 1
  53. 53
    Coding: Print Recall Score instead of Accuracy Score.
    Times asked 1
  54. 54
    Coding: Implement RandomForestRegressor (import, initialize, fit, predict).
    Times asked 1
  55. 55
    Coding: Build a simple preprocessing pipeline.
    Times asked 1
  56. 56
    Coding: Write a custom preprocessing pipeline.
    Times asked 1
  57. 57
    Coding: Apply Label Encoding to ["green", "blue", "white", "blue", "green"].
    Times asked 1
  58. 58
    Coding: Filter rows where RegionCode = "Florida".
    Times asked 1
  59. 59
    Coding: Filter rows where RegionCode = "Florida" and TargetValue > 50000.
    Times asked 1
  60. 60
    Coding: Create a subset of the data for RegionCode = "Florida".
    Times asked 1
  61. 61
    Coding: Drop two specified columns from the training dataset.
    Times asked 1
  1. 1
    Introduction
    Times asked 5
  2. 2
    Start explaining your notebook from the EDA section.
    Times asked 1
  3. 3
    Explain every line of your code in detail.
    Times asked 1
  4. 4
    No additional theory questions were asked.
    Times asked 2
  1. 1
    Introduction.
    Times asked 9
  2. 2
    Explain the problem statement.
    Times asked 6
  3. 3
    Explain your notebook from start to finish.
    Times asked 1
  4. 4
    What is the learning rate?
    Times asked 3
  5. 5
    How does the learning rate affect overfitting and underfitting?
    Times asked 1
  6. 6
    Coding: Create a dummy Pipeline and add it to a ColumnTransformer.
    Times asked 1
  7. 7
    Coding: Use a different scaling method and a different imputation strategy.
    Times asked 1
  8. 8
    What is Data Preprocessing?
    Times asked 1
  9. 9
    What is EDA?
    Times asked 1
  10. 10
    What is Feature Engineering?
    Times asked 1
  11. 11
    What is Hyperparameter Tuning?
    Times asked 4
  12. 12
    What does random_state = 42 mean?
    Times asked 1
  13. 13
    Questions based on your notebook implementation.
    Times asked 2
  1. 1
    Explain your notebook/project.
    Times asked 21
  2. 2
    How did you handle outliers?
    Times asked 3
  3. 3
    Is removing outliers always a good practice?
    Times asked 1
  4. 4
    What can you do instead of removing outliers?
    Times asked 1
  5. 5
    What is collinearity?
    Times asked 1
  6. 6
    If two features are highly correlated, what should you do?
    Times asked 1
  7. 7
    Explain Bagging.
    Times asked 6
  8. 8
    Explain Boosting.
    Times asked 6
  9. 9
    How does Bagging work?
    Times asked 1
  10. 10
    How does Boosting work?
    Times asked 1
  11. 11
    Explain Decision Tree.
    Times asked 3
  12. 12
    How does a Decision Tree work?
    Times asked 2
  13. 13
    What is Entropy?
    Times asked 2
  14. 14
    Why are Bagging and Boosting better than a single Decision Tree?
    Times asked 1
  15. 15
    What are different evaluation metrics?
    Times asked 1
  16. 16
    Write the formulas for evaluation metrics.
    Times asked 1
  17. 17
    Why is Accuracy not always a good evaluation metric?
    Times asked 1
  18. 18
    For a diabetes dataset, which evaluation metric would you choose and why?
    Times asked 1
  19. 19
    How does Logistic Regression work?
    Times asked 3
  20. 20
    How does KNN work?
    Times asked 1
  1. 1
    What is the problem statement?
    Times asked 1
  2. 2
    Explain your notebook/project.
    Times asked 21
  3. 3
    Explain Precision.
    Times asked 1
  4. 4
    Explain Recall.
    Times asked 1
  5. 5
    Explain F1-Score.
    Times asked 2
  6. 6
    Why did you use Label Encoding instead of other encoding techniques?
    Times asked 1
  7. 7
    Explain the hyperparameters you used.
    Times asked 1
  8. 8
    Explain Logistic Regression.
    Times asked 5
  9. 9
    Explain SVM.
    Times asked 3
  10. 10
    What is a kernel in SVM?
    Times asked 1
  11. 11
    Explain KNN.
    Times asked 3
  12. 12
    Why is KNN called a lazy learner?
    Times asked 2
  13. 13
    What are the different types of SVM kernels?
    Times asked 1
  14. 14
    Explain Bagging.
    Times asked 6
  15. 15
    Explain Boosting.
    Times asked 6
  16. 16
    Why did you use Random Forest?
    Times asked 2
  17. 17
    What are feature selection techniques?
    Times asked 2
  18. 18
    What is ROC-AUC?
    Times asked 1
  19. 19
    Why didn't you use ROC-AUC as the scoring metric?
    Times asked 1
  1. 1
    What are the techniques for handling null values?
    Times asked 1
  2. 2
    How do you handle imbalanced datasets?
    Times asked 2
  3. 3
    Why did you try multiple models?
    Times asked 2
  4. 4
    What is the relationship between Information Gain and Entropy?
    Times asked 1
  5. 5
    Explain the Decision Tree building process.
    Times asked 2
  6. 6
    Why did you choose that learning rate for LightGBM?
    Times asked 1
  7. 7
    Coding: On the MNIST dataset, create a Pipeline with StandardScaler and SVM, print the Classification Report, then apply PCA with SVM and compare the scores.
    Times asked 1
  8. 8
    Name different activation functions.
    Times asked 2
  9. 9
    Explain the Tanh activation function (input and output).
    Times asked 1
Maintaine By Lazy IITians Team + IITM BS students Regualr for More data you have