Skip to content

Commit 71b06ef

Browse files
committed
Deployed ef8f205 with MkDocs version: 1.6.1
1 parent 0f1e187 commit 71b06ef

6 files changed

Lines changed: 44 additions & 52 deletions

File tree

citing/index.html

Lines changed: 9 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -93,23 +93,28 @@
9393
<div class="section" itemprop="articleBody">
9494

9595
<!-- when new paper is published, add that paper and say "if you use SWRF, MultiSWRF/MultiSWRF*, or MultiSWRFDB/MultiSWRFDB*, cite the following paper" -->
96-
<p>If you use <strong>scikit-rebate</strong> or the <strong>MultiSURF</strong> algorithm in a scientific publication, please consider citing the following paper (currently available as a pre-print in arXiv):</p>
96+
<p>If you use <strong>scikit-rebate</strong> or the <strong>MultiSURF</strong> algorithm in a scientific publication, please consider citing the following paper:</p>
9797
<!-- *Urbanowicz, Ryan J., Randal S. Olson, Peter Schmitt, Melissa Meeker, and Jason H. Moore. "Benchmarking relief-based feature selection methods." arXiv preprint arXiv:1711.08477 (2017).* -->
9898
<p><em>Urbanowicz, Ryan J., Randal S. Olson, Peter Schmitt, Melissa Meeker, and Jason H. Moore. "Benchmarking relief-based feature selection methods for bioinformatics data mining." Journal of Biomedical Informatics, 85:168–188, 2018.</em></p>
9999
<p>Alternatively a complete <strong>review of Relief-based algorithms</strong> is available at:</p>
100-
<p><em>Urbanowicz, Ryan J., Melissa Meeker, William LaCava, Randal S. Olson, and Jason H. Moore. "Relief-based feature selection: introduction and review." arXiv preprint arXiv:1711.08421 (2017).</em></p>
100+
<!-- *Urbanowicz, Ryan J., Melissa Meeker, William LaCava, Randal S. Olson, and Jason H. Moore. "Relief-based feature selection: introduction and review." arXiv preprint arXiv:1711.08421 (2017).* -->
101+
<p><em>Urbanowicz, Ryan J., Melissa Meeker, William LaCava, Randal S. Olson, and Jason H. Moore. "Relief-based feature selection: introduction and review." Journal of Biomedical
102+
Informatics, 85:189–203, 2018.</em></p>
101103
<p>To cite the <strong>original Relief</strong> paper:</p>
102104
<p><em>Kira, Kenji, and Larry A. Rendell. "A practical approach to feature selection." In Machine Learning Proceedings 1992, pp. 249-256. 1992.</em></p>
103105
<p>To cite the <strong>original ReliefF</strong> paper: </p>
104106
<p><em>Kononenko, Igor. "Estimating attributes: analysis and extensions of RELIEF." In European conference on machine learning, pp. 171-182. Springer, Berlin, Heidelberg, 1994.</em></p>
105107
<p>To cite the <strong>original SURF</strong> paper:</p>
106-
<p><em>Greene, Casey S., Nadia M. Penrod, Jeff Kiralis, and Jason H. Moore. "Spatially uniform relieff (SURF) for computationally-efficient filtering of gene-gene interactions." BioData mining 2, no. 1 (2009): 5.</em></p>
108+
<p><em>Greene, Casey S., Nadia M. Penrod, Jeff Kiralis, and Jason H. Moore. "Spatially uniform relieff (SURF) for computationally-efficient filtering of gene-gene interactions." BioData Mining 2, no. 1 (2009): 5.</em></p>
107109
<p>To cite the <strong>original SURF*</strong> paper: </p>
108110
<p><em>Greene, Casey S., Daniel S. Himmelstein, Jeff Kiralis, and Jason H. Moore. "The informative extremes: using both nearest and farthest individuals can improve relief algorithms in the domain of human genetics." In European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics, pp. 182-193. Springer, Berlin, Heidelberg, 2010.</em></p>
109111
<p>To cite the <strong>original MultiSURF*</strong> paper:</p>
110112
<p><em>Granizo-Mackenzie, Delaney, and Jason H. Moore. "Multiple threshold spatially uniform relieff for the genetic analysis of complex human diseases." In European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics, pp. 1-10. Springer, Berlin, Heidelberg, 2013.</em></p>
111113
<!-- Add citations for SWRF* and μ-Relief below -->
112-
114+
<p>To cite the <strong>original SWRF*</strong> paper:</p>
115+
<p><em>Stokes, Matthew E., and Shyam Visweswaran. "Application of a spatially-weighted relief algorithm for ranking genetic predictors of disease." BioData Mining, 5:20, 2012.</em></p>
116+
<p>To cite the <strong>original μ-Relief</strong> paper:</p>
117+
<p><em>Aggarwal, Nitisha, Unmesh Shukla, G. J. Saxena, et al. "Mean based relief: An improved feature selection method based on relieff." Applied Intelligence, 53:23004–23028, 2023.</em></p>
113118
<p>To cite the <strong>original TuRF</strong> paper: </p>
114119
<p><em>Moore, Jason H., and Bill C. White. "Tuning ReliefF for genome-wide genetic analysis." In European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics, pp. 166-175. Springer, Berlin, Heidelberg, 2007.</em></p>
115120

index.html

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -168,5 +168,5 @@
168168

169169
<!--
170170
MkDocs version : 1.6.1
171-
Build Date UTC : 2026-04-25 00:32:34.114289+00:00
171+
Build Date UTC : 2026-04-27 22:30:18.205766+00:00
172172
-->

search/search_index.json

Lines changed: 1 addition & 1 deletion
Large diffs are not rendered by default.

sitemap.xml

Lines changed: 7 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -2,30 +2,30 @@
22
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
33
<url>
44
<loc>http://UrbsLab.github.io/scikit-rebate/</loc>
5-
<lastmod>2026-04-25</lastmod>
5+
<lastmod>2026-04-27</lastmod>
66
</url>
77
<url>
88
<loc>http://UrbsLab.github.io/scikit-rebate/citing/</loc>
9-
<lastmod>2026-04-25</lastmod>
9+
<lastmod>2026-04-27</lastmod>
1010
</url>
1111
<url>
1212
<loc>http://UrbsLab.github.io/scikit-rebate/contributing/</loc>
13-
<lastmod>2026-04-25</lastmod>
13+
<lastmod>2026-04-27</lastmod>
1414
</url>
1515
<url>
1616
<loc>http://UrbsLab.github.io/scikit-rebate/installing/</loc>
17-
<lastmod>2026-04-25</lastmod>
17+
<lastmod>2026-04-27</lastmod>
1818
</url>
1919
<url>
2020
<loc>http://UrbsLab.github.io/scikit-rebate/releases/</loc>
21-
<lastmod>2026-04-25</lastmod>
21+
<lastmod>2026-04-27</lastmod>
2222
</url>
2323
<url>
2424
<loc>http://UrbsLab.github.io/scikit-rebate/support/</loc>
25-
<lastmod>2026-04-25</lastmod>
25+
<lastmod>2026-04-27</lastmod>
2626
</url>
2727
<url>
2828
<loc>http://UrbsLab.github.io/scikit-rebate/using/</loc>
29-
<lastmod>2026-04-25</lastmod>
29+
<lastmod>2026-04-27</lastmod>
3030
</url>
3131
</urlset>

sitemap.xml.gz

0 Bytes
Binary file not shown.

using/index.html

Lines changed: 26 additions & 39 deletions
Original file line numberDiff line numberDiff line change
@@ -148,15 +148,16 @@ <h1 id="using-skrebate">Using skrebate</h1>
148148
<p>We have designed the Relief-based algorithms to be integrated directly into scikit-learn machine learning workflows. Below, we provide code samples showing how the various Relief-based algorithms can be used as feature selection methods in scikit-learn pipelines.</p>
149149
<p>For details on the algorithmic differences between the various Relief-based algorithms, please refer to <a href="https://arxiv.org/abs/1711.08477">this research paper</a>.</p>
150150
<h2 id="using-the-core-algorithms">Using the Core Algorithms</h2>
151-
<p>Core Relief-based algorithms are Relief-based algorithms that perform a single pass over the training data (i.e. each target instance is used one time). </p>
152-
<p>ReliefF was the original, most widely-known core Relief-based algorithm and it allows you to specify the number of nearest neighbors to consider during feature scoring.</p>
153-
<p>SURF, SURF*, MultiSURF, MultiSURF*, SWRF, SWRF*, MultiSWRF, MultiSWRF*, MultiSWRFDB, and MultiSWRFDB* are all extensions to the ReliefF algorithm that automatically determine the number of neighbors to consider when scoring the features. μ-Relief is an extension that, like ReliefF, requires a preset number of neighbors, but determines neighborhood membership differently than ReliefF.</p>
154-
<p>The hyperparameter settings and usage examples for each of these algorithms are provided below. </p>
155-
<h3 id="relieff">ReliefF<sup id="fnref:1"><a class="footnote-ref" href="#fn:1">1</a></sup></h3>
151+
<p>Core Relief-based algorithms are Relief-based algorithms (RBAs) that perform a single pass over the training data (i.e. each target instance is used once). </p>
152+
<p>ReliefF was the original, most widely-known core RBA and it allows you to specify the number of nearest neighbors to consider during feature scoring.</p>
153+
<p>SURF, SURF*, MultiSURF, MultiSURF*, SWRF, SWRF*, MultiSWRF, MultiSWRF*, MultiSWRFDB, and MultiSWRFDB* are all core RBAs that automatically determine the number of neighbors to consider when scoring the features. μ-Relief is a core RBA that, like ReliefF, requires a preset number of neighbors, but determines neighborhood membership differently than ReliefF.</p>
154+
<p>The hyperparameter settings and usage examples for each of these algorithms in scikit-rebate are provided below. </p>
155+
<h3 id="relieff">ReliefF</h3>
156156
<!-- ReliefF is the most basic of the Relief-based feature selection algorithms, and the implementation allows you to specify the number of nearest neighbors to consider in the scoring algorithm. The parameters for the ReliefF algorithm are as follows: -->
157157
<p>Determines neighborhood membership based on k-nearest neighbors (<code>n_neighbors</code>).</p>
158158
<!-- Includes k nearest instances as neighbors (`n_neighbors`). -->
159159

160+
<p>To learn more about the algorithm, read this <a href="https://doi.org/10.1007/3-540-57868-4_57">paper</a>.</p>
160161
<table>
161162
<thead>
162163
<tr>
@@ -277,9 +278,10 @@ <h3 id="relieff">ReliefF<sup id="fnref:1"><a class="footnote-ref" href="#fn:1">1
277278
278279
SURF, SURF\*, MultiSURF, MultiSURF\*, SWRF, SWRF\*, MultiSWRF, MultiSWRF\*, MultiSWRFDB, and MultiSWRFDB\* are all extensions to the ReliefF algorithm that automatically determine the ideal number of neighbors to consider when scoring the features. Note that all of these algorithms utilize the same group of hyperparameters, which are the same hyperparameters as ReliefF excluding `n_neighbors`. -->
279280

280-
<h3 id="surf">SURF<sup id="fnref:2"><a class="footnote-ref" href="#fn:2">2</a></sup></h3>
281+
<h3 id="surf">SURF</h3>
281282
<p>Includes instances closer than the global mean distance as near neighbors.</p>
282283
<p>The global mean distance is the mean of all pairwise distances in the dataset.</p>
284+
<p>To learn more about the algorithm, read this <a href="https://doi.org/10.1186/1756-0381-2-5">paper</a>.</p>
283285
<table>
284286
<thead>
285287
<tr>
@@ -388,9 +390,10 @@ <h3 id="surf">SURF<sup id="fnref:2"><a class="footnote-ref" href="#fn:2">2</a></
388390
&gt;&gt;&gt; N6 -0.00590625 19
389391
&gt;&gt;&gt; N3 -0.00634880 20
390392
</code></pre>
391-
<h3 id="surf_1">SURF*<sup id="fnref:3"><a class="footnote-ref" href="#fn:3">3</a></sup></h3>
393+
<h3 id="surf_1">SURF*</h3>
392394
<p>Includes instances closer than the global mean distance as near neighbors and instances farther than the global mean distance as far neighbors.</p>
393395
<p>The global mean distance is the mean of all pairwise distances in the dataset.</p>
396+
<p>To learn more about the algorithm, read this <a href="https://doi.org/10.1007/978-3-642-12211-8_16">paper</a>.</p>
394397
<table>
395398
<thead>
396399
<tr>
@@ -499,9 +502,10 @@ <h3 id="surf_1">SURF*<sup id="fnref:3"><a class="footnote-ref" href="#fn:3">3</a
499502
&gt;&gt;&gt; N6 -0.01174531 19
500503
&gt;&gt;&gt; N3 -0.01225238 20
501504
</code></pre>
502-
<h3 id="multisurf">MultiSURF<sup id="fnref:4"><a class="footnote-ref" href="#fn:4">4</a></sup></h3>
505+
<h3 id="multisurf">MultiSURF</h3>
503506
<p>Includes instances closer than <code>μ-σ/2</code> as near neighbors.</p>
504507
<p>Recomputes the mean distance and standard deviation per target instance (i.e. mean distance to the target instance and standard deviation of these distances).</p>
508+
<p>To learn more about the algorithm, read this <a href="https://doi.org/10.1016/j.jbi.2018.07.015">paper</a>.</p>
505509
<table>
506510
<thead>
507511
<tr>
@@ -610,9 +614,10 @@ <h3 id="multisurf">MultiSURF<sup id="fnref:4"><a class="footnote-ref" href="#fn:
610614
&gt;&gt;&gt; N3 -0.00931820 19
611615
&gt;&gt;&gt; N17 -0.00978209 20
612616
</code></pre>
613-
<h3 id="multisurf_1">MultiSURF*<sup id="fnref:5"><a class="footnote-ref" href="#fn:5">5</a></sup></h3>
617+
<h3 id="multisurf_1">MultiSURF*</h3>
614618
<p>Includes instances closer than <code>μ-σ/2</code> as near neighbors and instances farther than <code>μ+σ/2</code> as far neighbors. Instances within half a standard deviation of the mean distance are excluded from the neighborhood (are in the "deadband zone").</p>
615619
<p>Recomputes the mean distance and standard deviation per target instance (i.e. mean distance to the target instance and standard deviation of these distances).</p>
620+
<p>To learn more about the algorithm, read this <a href="https://doi.org/10.1007/978-3-642-37189-9_1">paper</a>.</p>
616621
<table>
617622
<thead>
618623
<tr>
@@ -832,9 +837,10 @@ <h3 id="swrf">SWRF</h3>
832837
&gt;&gt;&gt; N17 -0.01826211 19
833838
&gt;&gt;&gt; N6 -0.01848672 20
834839
</code></pre>
835-
<h3 id="swrf_1">SWRF*<sup id="fnref:6"><a class="footnote-ref" href="#fn:6">6</a></sup></h3>
840+
<h3 id="swrf_1">SWRF*</h3>
836841
<p>Adjusts weights given to neighbors through a sigmoidal gradient function. Weights decrease as distance to the target instance increases, and all instances are included in the neighborhood (instances farther than the global mean distance are added as far neighbors).</p>
837842
<p>The global mean distance is the mean of all pairwise distances in the dataset and the global standard deviation, used in the gradient function, is the standard deviation of these distances.</p>
843+
<p>To learn more about the algorithm, read this <a href="https://doi.org/10.1186/1756-0381-5-20">paper</a>.</p>
838844
<table>
839845
<thead>
840846
<tr>
@@ -1391,8 +1397,9 @@ <h3 id="multiswrfdb_1">MultiSWRFDB*</h3>
13911397
13921398
μ-Relief, like ReliefF, utilizes the `n_neighbors` hyperparameter. It has the same hyperparameters as ReliefF. -->
13931399

1394-
<h3 id="-relief">μ-Relief<sup id="fnref:7"><a class="footnote-ref" href="#fn:7">7</a></sup></h3>
1395-
<p>Includes as neighbors: the k (<code>n_neighbors</code>) instances whose absolute difference between 1) their distance from the target instance and 2) the mean distance among their class from the target instance is the greatest. </p>
1400+
<h3 id="-relief">μ-Relief</h3>
1401+
<p>Includes as neighbors: the k (<code>n_neighbors</code>) instances whose absolute difference between 1) their distance from the target instance and 2) the mean distance among their class from the target instance is the greatest.</p>
1402+
<p>To learn more about the algorithm, read this <a href="https://doi.org/10.1007/s10489-023-04662-w">paper</a>. </p>
13961403
<table>
13971404
<thead>
13981405
<tr>
@@ -1874,33 +1881,13 @@ <h2 id="general-usage-guidelines">General Usage Guidelines</h2>
18741881
<p>2.) In very large feature spaces users can expect core Relief-based algorithm scores to become less reliable when run on their own. This is because as the feature space becomes very large, the determination of nearest neighbors becomes more random. As a result, in very large feature spaces (e.g. &gt; 10,000 features), users should consider combining a core Relief-based algorithm with an iterative approach such as TuRF (implemented here), VLSRelief, or Iterative Relief. </p>
18751882
<p>3.) When scaling up to big data problems, keep in mind that the data aspect that slows down ReBATE methods the most is the number of training instances, since Relief-based algorithms scale linearly with the number of features, but quadratically with the number of training instances. This is the result of Relief-based methods needing to calculate a distance array (i.e. all pairwise distances between instances in the training dataset). If you have a very large number of training instances available, consider utilizing a class balanced random sampling of that dataset when running any ReBATE method to save on memory and computation time.</p>
18761883
<!-- References -->
1877-
<!-- [^4]: *Urbanowicz, Ryan J., Randal S. Olson, Peter Schmitt, Melissa Meeker, and Jason H. Moore. "Benchmarking relief-based feature selection methods." arXiv preprint arXiv:1711.08477 (2017).* -->
1878-
<div class="footnote">
1879-
<hr />
1880-
<ol>
1881-
<li id="fn:1">
1882-
<p><em>Kononenko, Igor. "Estimating attributes: analysis and extensions of RELIEF." In European conference on machine learning, pp. 171-182. Springer, Berlin, Heidelberg, 1994.</em>&#160;<a class="footnote-backref" href="#fnref:1" title="Jump back to footnote 1 in the text">&#8617;</a></p>
1883-
</li>
1884-
<li id="fn:2">
1885-
<p><em>Greene, Casey S., Nadia M. Penrod, Jeff Kiralis, and Jason H. Moore. "Spatially uniform relieff (SURF) for computationally-efficient filtering of gene-gene interactions." BioData mining 2, no. 1 (2009): 5.</em>&#160;<a class="footnote-backref" href="#fnref:2" title="Jump back to footnote 2 in the text">&#8617;</a></p>
1886-
</li>
1887-
<li id="fn:3">
1888-
<p><em>Greene, Casey S., Daniel S. Himmelstein, Jeff Kiralis, and Jason H. Moore. "The informative extremes: using both nearest and farthest individuals can improve relief algorithms in the domain of human genetics." In European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics, pp. 182-193. Springer, Berlin, Heidelberg, 2010.</em>&#160;<a class="footnote-backref" href="#fnref:3" title="Jump back to footnote 3 in the text">&#8617;</a></p>
1889-
</li>
1890-
<li id="fn:4">
1891-
<p><em>Urbanowicz, Ryan J., Randal S. Olson, Peter Schmitt, Melissa Meeker, and Jason H. Moore. "Benchmarking relief-based feature selection methods for bioinformatics data mining." Journal of Biomedical Informatics, 85:168–188, 2018.</em>&#160;<a class="footnote-backref" href="#fnref:4" title="Jump back to footnote 4 in the text">&#8617;</a></p>
1892-
</li>
1893-
<li id="fn:5">
1894-
<p><em>Granizo-Mackenzie, Delaney, and Jason H. Moore. "Multiple threshold spatially uniform relieff for the genetic analysis of complex human diseases." In European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics, pp. 1-10. Springer, Berlin, Heidelberg, 2013.</em>&#160;<a class="footnote-backref" href="#fnref:5" title="Jump back to footnote 5 in the text">&#8617;</a></p>
1895-
</li>
1896-
<li id="fn:6">
1897-
<p><em>Stokes, Matthew E., and Shyam Visweswaran. "Application of a spatially-weighted relief algorithm for ranking genetic predictors of disease." BioData Mining, 5:20, 2012.</em>&#160;<a class="footnote-backref" href="#fnref:6" title="Jump back to footnote 6 in the text">&#8617;</a></p>
1898-
</li>
1899-
<li id="fn:7">
1900-
<p>Aggarwal, Nitisha, Unmesh Shukla, G. J. Saxena, et al. Mean based relief: An improved feature selection method based on relieff. Applied Intelligence, 53:23004–23028, 2023. doi: 10.1007/s10489-023-04662-w.&#160;<a class="footnote-backref" href="#fnref:7" title="Jump back to footnote 7 in the text">&#8617;</a></p>
1901-
</li>
1902-
</ol>
1903-
</div>
1884+
<!-- [^1]: *Kononenko, Igor. "Estimating attributes: analysis and extensions of RELIEF." In European conference on machine learning, pp. 171-182. Springer, Berlin, Heidelberg, 1994.*
1885+
[^2]: *Greene, Casey S., Nadia M. Penrod, Jeff Kiralis, and Jason H. Moore. "Spatially uniform relieff (SURF) for computationally-efficient filtering of gene-gene interactions." BioData mining 2, no. 1 (2009): 5.*
1886+
[^3]: *Greene, Casey S., Daniel S. Himmelstein, Jeff Kiralis, and Jason H. Moore. "The informative extremes: using both nearest and farthest individuals can improve relief algorithms in the domain of human genetics." In European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics, pp. 182-193. Springer, Berlin, Heidelberg, 2010.*
1887+
[^4]: *Urbanowicz, Ryan J., Randal S. Olson, Peter Schmitt, Melissa Meeker, and Jason H. Moore. "Benchmarking relief-based feature selection methods for bioinformatics data mining." Journal of Biomedical Informatics, 85:168–188, 2018.*
1888+
[^5]: *Granizo-Mackenzie, Delaney, and Jason H. Moore. "Multiple threshold spatially uniform relieff for the genetic analysis of complex human diseases." In European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics, pp. 1-10. Springer, Berlin, Heidelberg, 2013.*
1889+
[^6]: *Stokes, Matthew E., and Shyam Visweswaran. "Application of a spatially-weighted relief algorithm for ranking genetic predictors of disease." BioData Mining, 5:20, 2012.*
1890+
[^7]: *Aggarwal, Nitisha, Unmesh Shukla, G. J. Saxena, et al. "Mean based relief: An improved feature selection method based on relieff." Applied Intelligence, 53:23004–23028, 2023. doi: 10.1007/s10489-023-04662-w.* -->
19041891

19051892
</div>
19061893
</div><footer>

0 commit comments

Comments
 (0)