Commits · a7317ccdc20d001e5b7f5277b0535923468bfbc6 · cs525-sp18-g07 / spark

Aug 14, 2015

[SPARK-8744] [ML] Add a public constructor to StringIndexer · a7317ccd

Holden Karau authored 9 years ago

It would be helpful to allow users to pass a pre-computed index to create an indexer, rather than always going through StringIndexer to create the model.

Author: Holden Karau <holden@pigscanfly.ca>

Closes #7267 from holdenk/SPARK-8744-StringIndexerModel-should-have-public-constructor.

a7317ccd

[SPARK-9956] [ML] Make trees work with one-category features · 7ecf0c46

Joseph K. Bradley authored 9 years ago

This modifies DecisionTreeMetadata construction to treat 1-category features as continuous, so that trees do not fail with such features. It is important for the pipelines API, where VectorIndexer can automatically categorize certain features as categorical.

As stated in the JIRA, this is a temp fix which we can improve upon later by automatically filtering out those features. That will take longer, though, since it will require careful indexing.

Targeted for 1.5 and master

CC: manishamde mengxr yanboliang

Author: Joseph K. Bradley <joseph@databricks.com>

Closes #8187 from jkbradley/tree-1cat.

7ecf0c46

[SPARK-9661] [MLLIB] minor clean-up of SPARK-9661 · a0e1abbd

Xiangrui Meng authored 9 years ago

Some minor clean-ups after SPARK-9661. See my inline comments. MechCoder jkbradley

Author: Xiangrui Meng <meng@databricks.com>

Closes #8190 from mengxr/SPARK-9661-fix.

a0e1abbd

[SPARK-9958] [SQL] Make HiveThriftServer2Listener thread-safe and update the... · c8677d73

zsxwing authored 9 years ago

[SPARK-9958] [SQL] Make HiveThriftServer2Listener thread-safe and update the tab name to "JDBC/ODBC Server"

This PR fixed the thread-safe issue of HiveThriftServer2Listener, and also changed the tab name to "JDBC/ODBC Server" since it's conflict with the new SQL tab.

<img width="1377" alt="thriftserver" src="https://cloud.githubusercontent.com/assets/1000778/9265707/c46f3f2c-4269-11e5-8d7e-888c9113ab4f.png">

Author: zsxwing <zsxwing@gmail.com>

Closes #8185 from zsxwing/SPARK-9958.

c8677d73

[MINOR] [SQL] Remove canEqual in Row · 7c7c7529

Liang-Chi Hsieh authored 9 years ago

As `InternalRow` does not extend `Row` now, I think we can remove it.

Author: Liang-Chi Hsieh <viirya@appier.com>

Closes #8170 from viirya/remove_canequal.

7c7c7529

Aug 13, 2015

[SPARK-9945] [SQL] pageSize should be calculated from executor.memory · bd35385d

Davies Liu authored 9 years ago

Currently, pageSize of TungstenSort is calculated from driver.memory, it should use executor.memory instead.

Also, in the worst case, the safeFactor could be 4 (because of rounding), increase it to 16.

cc rxin

Author: Davies Liu <davies@databricks.com>

Closes #8175 from davies/page_size.

bd35385d

[SPARK-9580] [SQL] Replace singletons in SQL tests · 8187b3ae

Andrew Or authored 9 years ago

A fundamental limitation of the existing SQL tests is that *there is simply no way to create your own `SparkContext`*. This is a serious limitation because the user may wish to use a different master or config. As a case in point, `BroadcastJoinSuite` is entirely commented out because there is no way to make it pass with the existing infrastructure.

This patch removes the singletons `TestSQLContext` and `TestData`, and instead introduces a `SharedSQLContext` that starts a context per suite. Unfortunately the singletons were so ingrained in the SQL tests that this patch necessarily needed to touch *all* the SQL test files.

<!-- Reviewable:start -->
[<img src="https://reviewable.io/review_button.png" height=40 alt="Review on Reviewable"/>](https://reviewable.io/reviews/apache/spark/8111)
<!-- Reviewable:end -->

Author: Andrew Or <andrew@databricks.com>

Closes #8111 from andrewor14/sql-tests-refactor.

8187b3ae

[SPARK-9943] [SQL] deserialized UnsafeHashedRelation should be serializable · c50f97da

Davies Liu authored 9 years ago

When the free memory in executor goes low, the cached broadcast objects need to serialized into disk, but currently the deserialized UnsafeHashedRelation can't be serialized , fail with NPE. This PR fixes that.

cc rxin

Author: Davies Liu <davies@databricks.com>

Closes #8174 from davies/serialize_hashed.

c50f97da

[SPARK-8976] [PYSPARK] fix open mode in python3 · 693949ba

Davies Liu authored 9 years ago

This bug only happen on Python 3 and Windows.

I tested this manually with python 3 and disable python daemon, no unit test yet.

Author: Davies Liu <davies@databricks.com>

Closes #8181 from davies/open_mode.

693949ba

[SPARK-9922] [ML] rename StringIndexerReverse to IndexToString · 6c5858bc

Xiangrui Meng authored 9 years ago

What `StringIndexerInverse` does is not strictly associated with `StringIndexer`, and the name is not clearly describing the transformation. Renaming to `IndexToString` might be better.

~~I also changed `invert` to `inverse` without arguments. `inputCol` and `outputCol` could be set after.~~
I also removed `invert`.

jkbradley holdenk

Author: Xiangrui Meng <meng@databricks.com>

Closes #8152 from mengxr/SPARK-9922.

6c5858bc

[SPARK-9935] [SQL] EqualNotNull not processed in ORC · c2520f50

hyukjinkwon authored 9 years ago

https://issues.apache.org/jira/browse/SPARK-9935

Author: hyukjinkwon <gurwls223@gmail.com>

Closes #8163 from HyukjinKwon/master.

c2520f50

[SPARK-9942] [PYSPARK] [SQL] ignore exceptions while try to import pandas · a8d2f4c5

Davies Liu authored 9 years ago

If pandas is broken (can't be imported, raise other exceptions other than ImportError), pyspark can't be imported, we should ignore all the exceptions.

Author: Davies Liu <davies@databricks.com>

Closes #8173 from davies/fix_pandas.

a8d2f4c5

[SPARK-9661] [MLLIB] [ML] Java compatibility · 864de8ea

MechCoder authored 9 years ago

I skimmed through the docs for various instance of Object and replaced them with Java compaible versions of the same.

1. Some methods in LDAModel.
2. runMiniBatchSGD
3. kolmogorovSmirnovTest

Author: MechCoder <manojkumarsivaraj334@gmail.com>

Closes #8126 from MechCoder/java_incop.

864de8ea

[SPARK-9649] Fix MasterSuite, third time's a charm · 8815ba2f

Andrew Or authored 9 years ago

This particular test did not load the default configurations so
it continued to start the REST server, which causes port bind
exceptions.

8815ba2f

[MINOR] [DOC] fix mllib pydoc warnings · 65fec798

Xiangrui Meng authored 9 years ago

Switch to correct Sphinx syntax. MechCoder

Author: Xiangrui Meng <meng@databricks.com>

Closes #8169 from mengxr/mllib-pydoc-fix.

65fec798

[MINOR] [ML] change MultilayerPerceptronClassifierModel to MultilayerPerceptronClassificationModel · 4b70798c

Yanbo Liang authored 9 years ago

To follow the naming rule of ML, change `MultilayerPerceptronClassifierModel` to `MultilayerPerceptronClassificationModel` like `DecisionTreeClassificationModel`, `GBTClassificationModel` and so on.

Author: Yanbo Liang <ybliang8@gmail.com>

Closes #8164 from yanboliang/mlp-name.

4b70798c

[SPARK-8965] [DOCS] Add ml-guide Python Example: Estimator, Transformer, and Param · 7a539ef3

Rosstin authored 9 years ago

Added ml-guide Python Example: Estimator, Transformer, and Param
/docs/_site/ml-guide.html

Author: Rosstin <asterazul@gmail.com>

Closes #8081 from Rosstin/SPARK-8965.

7a539ef3

[SPARK-9073] [ML] spark.ml Models copy() should call setParent when there is a parent · 2932e25d

lewuathe authored 9 years ago

Copied ML models must have the same parent of original ones

Author: lewuathe <lewuathe@me.com>
Author: Lewuathe <lewuathe@me.com>

Closes #7447 from Lewuathe/SPARK-9073.

2932e25d

[SPARK-9757] [SQL] Fixes persistence of Parquet relation with decimal column · 69930310

Cheng Lian authored 9 years ago

PR #7967 enables us to save data source relations to metastore in Hive compatible format when possible. But it fails to persist Parquet relations with decimal column(s) to Hive metastore of versions lower than 1.2.0. This is because `ParquetHiveSerDe` in Hive versions prior to 1.2.0 doesn't support decimal. This PR checks for this case and falls back to Spark SQL specific metastore table format.

Author: Yin Huai <yhuai@databricks.com>
Author: Cheng Lian <lian@databricks.com>

Closes #8130 from liancheng/spark-9757/old-hive-parquet-decimal.

69930310

[SPARK-9885] [SQL] Also pass barrierPrefixes and sharedPrefixes to... · 84a27916

Yin Huai authored 9 years ago

[SPARK-9885] [SQL] Also pass barrierPrefixes and sharedPrefixes to IsolatedClientLoader when hiveMetastoreJars is set to maven.

https://issues.apache.org/jira/browse/SPARK-9885

cc marmbrus liancheng

Author: Yin Huai <yhuai@databricks.com>

Closes #8158 from yhuai/classloaderMaven.

84a27916

[SPARK-9918] [MLLIB] remove runs from k-means and rename epsilon to tol · 68f99571

Xiangrui Meng authored 9 years ago

This requires some discussion. I'm not sure whether `runs` is a useful parameter. It certainly complicates the implementation. We might want to optimize the k-means implementation with block matrix operations. In this case, having `runs` may not be worth the trade-off. Also it increases the communication cost in a single job, which might cause other issues.

This PR also renames `epsilon` to `tol` to have consistent naming among algorithms. The Python constructor is updated to include all parameters.

jkbradley yu-iskw

Author: Xiangrui Meng <meng@databricks.com>

Closes #8148 from mengxr/SPARK-9918 and squashes the following commits:

149b9e5 [Xiangrui Meng] fix constructor in Python and rename epsilon to tol
3cc15b3 [Xiangrui Meng] fix test and change initStep to initSteps in python
a0a0274 [Xiangrui Meng] remove runs from k-means in the pipeline API

68f99571

[SPARK-9927] [SQL] Revert 8049 since it's pushing wrong filter down · d0b18919

Yijie Shen authored 9 years ago

I made a mistake in #8049 by casting literal value to attribute's data type, which would cause simply truncate the literal value and push a wrong filter down.

JIRA: https://issues.apache.org/jira/browse/SPARK-9927

Author: Yijie Shen <henry.yijieshen@gmail.com>

Closes #8157 from yjshen/rever8049.

d0b18919

[SPARK-9914] [ML] define setters explicitly for Java and use setParam group in RFormula · d7eb371e

Xiangrui Meng authored 9 years ago

The problem with defining setters in the base class is that it doesn't return the correct type in Java.

ericl

Author: Xiangrui Meng <meng@databricks.com>

Closes #8143 from mengxr/SPARK-9914 and squashes the following commits:

d36c887 [Xiangrui Meng] remove setters from model
a49021b [Xiangrui Meng] define setters explicitly for Java and use setParam group

d7eb371e

Aug 12, 2015

[SPARK-8922] [DOCUMENTATION, MLLIB] Add @since tags to mllib.evaluation · df543892
shikai.tang authored 9 years ago
```
Author: shikai.tang <tar.sky06@gmail.com>

Closes #7429 from mosessky/master.
```
df543892
[SPARK-9917] [ML] add getMin/getMax and doc for originalMin/origianlMax in MinMaxScaler · 5fc058a1
Xiangrui Meng authored 9 years ago
```
hhbyyh

Author: Xiangrui Meng <meng@databricks.com>

Closes #8145 from mengxr/SPARK-9917.
```
5fc058a1

[SPARK-9832] [SQL] add a thread-safe lookup for BytesToBytseMap · a8ab2634

Davies Liu authored 9 years ago

This patch add a thread-safe lookup for BytesToBytseMap, and use that in broadcasted HashedRelation.

Author: Davies Liu <davies@databricks.com>

Closes #8151 from davies/safeLookup.

a8ab2634

[SPARK-9920] [SQL] The simpleString of TungstenAggregate does not show its output · 22782190

Yin Huai authored 9 years ago

https://issues.apache.org/jira/browse/SPARK-9920

Taking `sqlContext.sql("select i, sum(j1) as sum from testAgg group by i").explain()` as an example, the output of our current master is
```
== Physical Plan ==
TungstenAggregate(key=[i#0], value=[(sum(cast(j1#1 as bigint)),mode=Final,isDistinct=false)]
 TungstenExchange hashpartitioning(i#0)
  TungstenAggregate(key=[i#0], value=[(sum(cast(j1#1 as bigint)),mode=Partial,isDistinct=false)]
   Scan ParquetRelation[file:/user/hive/warehouse/testagg][i#0,j1#1]
```
With this PR, the output will be
```
== Physical Plan ==
TungstenAggregate(key=[i#0], functions=[(sum(cast(j1#1 as bigint)),mode=Final,isDistinct=false)], output=[i#0,sum#18L])
 TungstenExchange hashpartitioning(i#0)
  TungstenAggregate(key=[i#0], functions=[(sum(cast(j1#1 as bigint)),mode=Partial,isDistinct=false)], output=[i#0,currentSum#22L])
   Scan ParquetRelation[file:/user/hive/warehouse/testagg][i#0,j1#1]
```

Author: Yin Huai <yhuai@databricks.com>

Closes #8150 from yhuai/SPARK-9920.

22782190

[SPARK-9916] [BUILD] [SPARKR] removed left-over sparkr.zip copy/create commands from codebase · 2fb4901b

Burak Yavuz authored 9 years ago

sparkr.zip is now built by SparkSubmit on a need-to-build basis.

cc shivaram

Author: Burak Yavuz <brkyvz@gmail.com>

Closes #8147 from brkyvz/make-dist-fix.

2fb4901b

[SPARK-9903] [MLLIB] skip local processing in PrefixSpan if there are no small prefixes · d7053bea

Xiangrui Meng authored 9 years ago

There exists a chance that the prefixes keep growing to the maximum pattern length. Then the final local processing step becomes unnecessary. feynmanliang

Author: Xiangrui Meng <meng@databricks.com>

Closes #8136 from mengxr/SPARK-9903.

d7053bea

[SPARK-9704] [ML] Made ProbabilisticClassifier, Identifiable, VectorUDT public APIs · d2d5e7fe

Joseph K. Bradley authored 9 years ago

Made ProbabilisticClassifier, Identifiable, VectorUDT public. All are annotated as DeveloperApi.

CC: mengxr EronWright

Author: Joseph K. Bradley <joseph@databricks.com>

Closes #8004 from jkbradley/ml-api-public-items and squashes the following commits:

7ebefda [Joseph K. Bradley] update per code review
7ff0768 [Joseph K. Bradley] attepting to add mima fix
756d84c [Joseph K. Bradley] VectorUDT annotated as AlphaComponent
ae7767d [Joseph K. Bradley] added another warning
94fd553 [Joseph K. Bradley] Made ProbabilisticClassifier, Identifiable, VectorUDT public APIs

d2d5e7fe

[SPARK-9908] [SQL] When spark.sql.tungsten.enabled is false, broadcast join does not work · 4413d085
Yin Huai authored 9 years ago
```
https://issues.apache.org/jira/browse/SPARK-9908

Author: Yin Huai <yhuai@databricks.com>

Closes #8149 from yhuai/SPARK-9908.
```
4413d085

[SPARK-9827] [SQL] fix fd leak in UnsafeRowSerializer · 7c35746c

Davies Liu authored 9 years ago

Currently, UnsafeRowSerializer does not close the InputStream, will cause fd leak if the InputStream has an open fd in it.

TODO: the fd could still be leaked, if any items in the stream is not consumed. Currently it replies on GC to close the fd in this case.

cc JoshRosen

Author: Davies Liu <davies@databricks.com>

Closes #8116 from davies/fd_leak.

7c35746c

[SPARK-9870] Disable driver UI and Master REST server in SparkSubmitSuite · 7b13ed27

Josh Rosen authored 9 years ago

I think that we should pass additional configuration flags to disable the driver UI and Master REST server in SparkSubmitSuite and HiveSparkSubmitSuite. This might cut down on port-contention-related flakiness in Jenkins.

Author: Josh Rosen <joshrosen@databricks.com>

Closes #8124 from JoshRosen/disable-ui-in-sparksubmitsuite.

7b13ed27

[SPARK-9855] [SPARKR] Add expression functions into SparkR whose params are simple · f4bc01f1

Yu ISHIKAWA authored 9 years ago

I added lots of expression functions for SparkR. This PR includes only functions whose params  are only `(Column)` or `(Column, Column)`.  And I think we need to improve how to test those functions. However, it would be better to work on another issue.

## Diff Summary

- Add lots of functions in `functions.R` and their generic in `generic.R`
- Add aliases for `ceiling` and `sign`
- Move expression functions from `column.R` to `functions.R`
- Modify `rdname` from `column` to `functions`

I haven't supported `not` function, because the name has a collesion with `testthat` package. I didn't think of the way  to define it.

## New Supported Functions

```
approxCountDistinct
ascii
base64
bin
bitwiseNOT
ceil (alias: ceiling)
crc32
dayofmonth
dayofyear
explode
factorial
hex
hour
initcap
isNaN
last_day
length
log2
ltrim
md5
minute
month
negate
quarter
reverse
round
rtrim
second
sha1
signum (alias: sign)
size
soundex
to_date
trim
unbase64
unhex
weekofyear
year

datediff
levenshtein
months_between
nanvl
pmod
```

## JIRA
[[SPARK-9855] Add expression functions into SparkR whose params are simple - ASF JIRA](https://issues.apache.org/jira/browse/SPARK-9855)

Author: Yu ISHIKAWA <yuu.ishikawa@gmail.com>

Closes #8123 from yu-iskw/SPARK-9855.

f4bc01f1

[SPARK-9724] [WEB UI] Avoid unnecessary redirects in the Spark Web UI. · 0d1d146c

Rohit Agarwal authored 9 years ago

Author: Rohit Agarwal <rohita@qubole.com>

Closes #8014 from mindprince/SPARK-9724 and squashes the following commits:

a7af5ff [Rohit Agarwal] [SPARK-9724] [WEB UI] Inline attachPrefix and attachPrefixForRedirect. Fix logic of attachPrefix
8a977cd [Rohit Agarwal] [SPARK-9724] [WEB UI] Address review comments: Remove unneeded code, update scaladoc.
b257844 [Rohit Agarwal] [SPARK-9724] [WEB UI] Avoid unnecessary redirects in the Spark Web UI.

0d1d146c

[SPARK-9780] [STREAMING] [KAFKA] prevent NPE if KafkaRDD instantiation … · 8ce60963

cody koeninger authored 9 years ago

…fails

Author: cody koeninger <cody@koeninger.org>

Closes #8133 from koeninger/SPARK-9780 and squashes the following commits:

406259d [cody koeninger] [SPARK-9780][Streaming][Kafka] prevent NPE if KafkaRDD instantiation fails

8ce60963

[SPARK-9449] [SQL] Include MetastoreRelation's inputFiles · 660e6dcf
Michael Armbrust authored 9 years ago
```
Author: Michael Armbrust <michael@databricks.com>

Closes #8119 from marmbrus/metastoreInputFiles.
```
660e6dcf
[SPARK-9915] [ML] stopWords should use StringArrayParam · fc1c7fd6
Xiangrui Meng authored 9 years ago
```
hhbyyh

Author: Xiangrui Meng <meng@databricks.com>

Closes #8141 from mengxr/SPARK-9915.
```
fc1c7fd6

[SPARK-9912] [MLLIB] QRDecomposition should use QType and RType for type names... · e6aef557

Xiangrui Meng authored 9 years ago

[SPARK-9912] [MLLIB] QRDecomposition should use QType and RType for type names instead of UType and VType

hhbyyh

Author: Xiangrui Meng <meng@databricks.com>

Closes #8140 from mengxr/SPARK-9912.

e6aef557

[SPARK-9909] [ML] [TRIVIAL] move weightCol to shared params · 6e409bc1

Holden Karau authored 9 years ago

As per the TODO move weightCol to Shared Params.

Author: Holden Karau <holden@pigscanfly.ca>

Closes #8144 from holdenk/SPARK-9909-move-weightCol-toSharedParams.

6e409bc1